Paper Page - Rex-Thinker: Grounded Object Referring Via Chain-of-Thought Reasoning

Rex-Thinker is a CoT-based model that enhances object referring by performing step-by-step reasoning over candidate objects, leading to improved interpretability and rejection of mismatched queries.

Object referring aims to detect all objects in an image that match a given
natural language description. We argue that a robust object referring model
should be grounded, meaning its predictions should be both explainable and
faithful to the visual content. Specifically, it should satisfy two key
properties: 1) Verifiable, by producing interpretable reasoning that justifies
its predictions and clearly links them to visual evidence; and 2) Trustworthy,
by learning to abstain when no object in the image satisfies the given
expression. However, most methods treat referring as a direct bounding box
prediction task, offering limited interpretability and struggling to reject
expressions with no matching object. In this work, we propose Rex-Thinker, a
model that formulates object referring as an explicit CoT reasoning task. Given
a referring expression, we first identify all candidate object instances
corresponding to the referred object category. Rex-Thinker then performs
step-by-step reasoning over each candidate to assess whether it matches the
given expression, before making a final prediction. To support this paradigm,
we construct a large-scale CoT-style referring dataset named HumanRef-CoT by
prompting GPT-4o on the HumanRef dataset. Each reasoning trace follows a
structured planning, action, and summarization format, enabling the model to
learn decomposed, interpretable reasoning over object candidates. We then train
Rex-Thinker in two stages: a cold-start supervised fine-tuning phase to teach
the model how to perform structured reasoning, followed by GRPO-based RL
learning to improve accuracy and generalization. Experiments show that our
approach outperforms standard baselines in both precision and interpretability
on in-domain evaluation, while also demonstrating improved ability to reject
hallucinated outputs and strong generalization in out-of-domain settings.

Source link

What's Hot

Perplexity AI tables $34.5 billion cash bid for Google’s Chrome amid Antitrust pressure

Moveworks Recognized as a Challenger in the 2025 Gartner® Magic Quadrant™ for Artificial Intelligence Applications in IT Service Management

Mobile and Voice Are Coming to Harvey – Artificial Lawyer

Paper page – Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding – Takara TLDR

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward – Takara TLDR

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions – Takara TLDR

Growing Support for Parthenon Marbles’ Return to Greece, More Art News

Leon Black and Leslie Wexner’s Letters to Jeffrey Epstein Released

School of Visual Arts Transfers Ownership to Nonprofit Alumni Society

Cristin Tierney Moves Gallery to Tribeca for 15th Anniversary Exhibition

Perplexity AI tables $34.5 billion cash bid for Google’s Chrome amid Antitrust pressure

Moveworks Recognized as a Challenger in the 2025 Gartner® Magic Quadrant™ for Artificial Intelligence Applications in IT Service Management

Mobile and Voice Are Coming to Harvey – Artificial Lawyer

What's Hot

Paper page – Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

Related Posts

Subscribe to Updates