Paper Page - UniVG-R1: Reasoning Guided Universal Visual Grounding With Reinforcement Learning

UniVG-R1, a reasoning-guided multimodal large language model, enhances visual grounding by leveraging reinforcement learning and a difficulty-aware strategy, achieving state-of-the-art results and strong generalizability.

Traditional visual grounding methods primarily focus on single-image
scenarios with simple textual references. However, extending these methods to
real-world scenarios that involve implicit and complex instructions,
particularly in conjunction with multiple images, poses significant challenges,
which is mainly due to the lack of advanced reasoning ability across diverse
multi-modal contexts. In this work, we aim to address the more practical
universal grounding task, and propose UniVG-R1, a reasoning guided multimodal
large language model (MLLM) for universal visual grounding, which enhances
reasoning capabilities through reinforcement learning (RL) combined with
cold-start data. Specifically, we first construct a high-quality
Chain-of-Thought (CoT) grounding dataset, annotated with detailed reasoning
chains, to guide the model towards correct reasoning paths via supervised
fine-tuning. Subsequently, we perform rule-based reinforcement learning to
encourage the model to identify correct reasoning chains, thereby incentivizing
its reasoning capabilities. In addition, we identify a difficulty bias arising
from the prevalence of easy samples as RL training progresses, and we propose a
difficulty-aware weight adjustment strategy to further strengthen the
performance. Experimental results demonstrate the effectiveness of UniVG-R1,
which achieves state-of-the-art performance on MIG-Bench with a 9.1%
improvement over the previous method. Furthermore, our model exhibits strong
generalizability, achieving an average improvement of 23.4% in zero-shot
performance across four image and video reasoning grounding benchmarks. The
project page can be accessed at https://amap-ml.github.io/UniVG-R1-page/.

Source link

What's Hot

We keep talking about AI agents, but do we ever know what they are?

Great Customer Service Will Be People And Bots Working Together

Samsung Galaxy S26 Rumors: Everything We Know So Far

Paper page – UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

PickStyle: Video-to-Video Style Transfer with Context-Style Adapters – Takara TLDR

OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment – Takara TLDR

GCPO: When Contrast Fails, Go Gold – Takara TLDR

Smithsonian Closes Museums Amid Government Shutdown

The Rubin Names 2025 Art Prize, Research and Art Projects Grants

Kochi-Muziris Biennial Announces 66 Artists for December Exhibition

Instagram Launches ‘Rings’ Awards for Creators—With KAWS as a Judge

We keep talking about AI agents, but do we ever know what they are?

Great Customer Service Will Be People And Bots Working Together

Samsung Galaxy S26 Rumors: Everything We Know So Far

What's Hot

Paper page – UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Related Posts

Subscribe to Updates