Paper Page - TinyLLaVA-Video-R1: Towards Smaller LMMs For Video Reasoning

Recently, improving the reasoning ability of large multimodal models (LMMs)
through reinforcement learning has made great progress. However, most existing
works are based on highly reasoning-intensive datasets such as mathematics and
code, and researchers generally choose large-scale models as the foundation. We
argue that exploring small-scale models’ reasoning capabilities remains
valuable for researchers with limited computational resources. Moreover,
enabling models to explain their reasoning processes on general
question-answering datasets is equally meaningful. Therefore, we present the
small-scale video reasoning model TinyLLaVA-Video-R1. Based on TinyLLaVA-Video,
a traceably trained video understanding model with no more than 4B parameters,
it not only demonstrates significantly improved reasoning and thinking
capabilities after using reinforcement learning on general Video-QA datasets,
but also exhibits the emergent characteristic of “aha moments”. Furthermore, we
share a series of experimental findings, aiming to provide practical insights
for future exploration of video reasoning (thinking) abilities in small-scale
models. It is available at https://github.com/ZhangXJ199/TinyLLaVA-Video-R1.

Source link

What's Hot

Legal Education Must Change Because of AI – Survey – Artificial Lawyer

BaseReward: A Strong Baseline for Multimodal Reward Model – Takara TLDR

Abu Dhabi’s TII and NVIDIA Launch Middle East’s First Joint ‘AI & Robotics’ NVAITC Research Lab

Paper page – TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

BaseReward: A Strong Baseline for Multimodal Reward Model – Takara TLDR

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer – Takara TLDR

RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation – Takara TLDR

New Collectors Drive Strong Sales at New York Fair

Hidden Portrait May Be Vermeer’s Earliest Known Work

Who Are the Art World Figures on the Time 100 List?

Acquavella Signs Harumi Klossowska de Rola, Daughter of Balthus

Legal Education Must Change Because of AI – Survey – Artificial Lawyer

BaseReward: A Strong Baseline for Multimodal Reward Model – Takara TLDR

Abu Dhabi’s TII and NVIDIA Launch Middle East’s First Joint ‘AI & Robotics’ NVAITC Research Lab

What's Hot

Paper page – TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Related Posts

Subscribe to Updates