Paper Page - PyVision: Agentic Vision With Dynamic Tooling

PyVision, an interactive framework, enables LLMs to autonomously create and refine Python-based tools for visual reasoning, achieving significant performance improvements across benchmarks.

LLMs are increasingly deployed as agents, systems capable of planning,
reasoning, and dynamically calling external tools. However, in visual
reasoning, prior approaches largely remain limited by predefined workflows and
static toolsets. In this report, we present PyVision, an interactive,
multi-turn framework that enables MLLMs to autonomously generate, execute, and
refine Python-based tools tailored to the task at hand, unlocking flexible and
interpretable problem-solving. We develop a taxonomy of the tools created by
PyVision and analyze their usage across a diverse set of benchmarks.
Quantitatively, PyVision achieves consistent performance gains, boosting
GPT-4.1 by +7.8% on V* and Claude-4.0-Sonnet by +31.1% on VLMsAreBlind-mini.
These results point to a broader shift: dynamic tooling allows models not just
to use tools, but to invent them, advancing toward more agentic visual
reasoning.

Source link

What's Hot

Has the Open Source Route Been Successful?_Points_model_has

MIT and Hasso Plattner Institute Unveil SustainaPrint, The Future of

Perplexity CEO Says Curiosity, Not Hype, Will Shape AI’s Future

Paper page – PyVision: Agentic Vision with Dynamic Tooling

Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth – Takara TLDR

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding – Takara TLDR

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers – Takara TLDR

Tony Shafrazi and the Art of the Comeback

Basquiats Linked to 1MDB Scandal Auctioned by US Government

US Ambassador to UK Fills Residence with Impressionist Masters

New Code of Ethics Implores UK Museums to End Fossil Fuel Sponsorships

Has the Open Source Route Been Successful?_Points_model_has

MIT and Hasso Plattner Institute Unveil SustainaPrint, The Future of

Perplexity CEO Says Curiosity, Not Hype, Will Shape AI’s Future

What's Hot

Paper page – PyVision: Agentic Vision with Dynamic Tooling

Related Posts

Subscribe to Updates