When And What: Diffusion-Grounded VideoLLM With Entity Aware Segmentation For Long Video Understanding - Takara TLDR

Understanding videos requires more than answering open ended questions, it
demands the ability to pinpoint when events occur and how entities interact
across time. While recent Video LLMs have achieved remarkable progress in
holistic reasoning, they remain coarse in temporal perception: timestamps are
encoded only implicitly, frame level features are weak in capturing continuity,
and language vision alignment often drifts from the entities of interest. In
this paper, we present Grounded VideoDiT, a Video LLM designed to overcome
these limitations by introducing three key innovations. First, a Diffusion
Temporal Latent (DTL) encoder enhances boundary sensitivity and maintains
temporal consistency. Second, object grounded representations explicitly bind
query entities to localized visual evidence, strengthening alignment. Third, a
mixed token scheme with discrete temporal tokens provides explicit timestamp
modeling, enabling fine grained temporal reasoning. Together, these designs
equip Grounded VideoDiT with robust grounding capabilities, as validated by
state of the art results on Charades STA, NExT GQA, and multiple VideoQA
benchmarks.

Source link

What's Hot

A&O Shearman Spin-Off aosphere Buys Investment Navigator – Updated – Artificial Lawyer

SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus – Takara TLDR

Indian Enterprises Put Key AI Roles in the Leadership Table: IBM Study

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding – Takara TLDR

SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus – Takara TLDR

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents – Takara TLDR

Improving GUI Grounding with Explicit Position-to-Coordinate Mapping – Takara TLDR

Sotheby’s to Sell René Magritte Held in Same Collection for 100 years

Former ARTnews Publisher Dies at 97

National Gallery of Art Closes as a Result of Government Shutdown

Almine Rech Closes London Gallery After More Than a Decade

A&O Shearman Spin-Off aosphere Buys Investment Navigator – Updated – Artificial Lawyer

SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus – Takara TLDR

Indian Enterprises Put Key AI Roles in the Leadership Table: IBM Study

What's Hot

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding – Takara TLDR

Related Posts

Subscribe to Updates