Deforming Videos To Masks: Flow Matching For Referring Video Segmentation - Takara TLDR

Referring Video Object Segmentation (RVOS) requires segmenting specific
objects in a video guided by a natural language description. The core challenge
of RVOS is to anchor abstract linguistic concepts onto a specific set of pixels
and continuously segment them through the complex dynamics of a video. Faced
with this difficulty, prior work has often decomposed the task into a pragmatic
`locate-then-segment’ pipeline. However, this cascaded design creates an
information bottleneck by simplifying semantics into coarse geometric prompts
(e.g, point), and struggles to maintain temporal consistency as the segmenting
process is often decoupled from the initial language grounding. To overcome
these fundamental limitations, we propose FlowRVS, a novel framework that
reconceptualizes RVOS as a conditional continuous flow problem. This allows us
to harness the inherent strengths of pretrained T2V models, fine-grained pixel
control, text-video semantic alignment, and temporal coherence. Instead of
conventional generating from noise to mask or directly predicting mask, we
reformulate the task by learning a direct, language-guided deformation from a
video’s holistic representation to its target mask. Our one-stage, generative
approach achieves new state-of-the-art results across all major RVOS
benchmarks. Specifically, achieving a $\mathcal{J}\&\mathcal{F}$ of 51.1 in
MeViS (+1.6 over prior SOTA) and 73.3 in the zero shot Ref-DAVIS17 (+2.7),
demonstrating the significant potential of modeling video understanding tasks
as continuous deformation processes.

Source link

What's Hot

Nexl Bags $23m, Will Invest In Hires + Acquisitions – Artificial Lawyer

ASPO: Asymmetric Importance Sampling Policy Optimization – Takara TLDR

Vxceed builds the perfect sales pitch for sales teams at scale using Amazon Bedrock

Deforming Videos to Masks: Flow Matching for Referring Video Segmentation – Takara TLDR

ASPO: Asymmetric Importance Sampling Policy Optimization – Takara TLDR

Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation – Takara TLDR

Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context – Takara TLDR

Matthiesen Gallery Files Lawsuit Over Gustave Courbet Painting

MoMA Partners with Mattel for Van Gogh Barbie, Monet and Dalí Figures

Underground Film Legend and Artist Dies at 92

Artwork Forfeited by Inigo Philbrick’s Partner Flops at Sotheby’s

Nexl Bags $23m, Will Invest In Hires + Acquisitions – Artificial Lawyer

ASPO: Asymmetric Importance Sampling Policy Optimization – Takara TLDR

Vxceed builds the perfect sales pitch for sales teams at scale using Amazon Bedrock

What's Hot

Deforming Videos to Masks: Flow Matching for Referring Video Segmentation – Takara TLDR

Related Posts

Subscribe to Updates