Paper page - PyVision: Agentic Vision with Dynamic Tooling

PyVision, an interactive framework, enables LLMs to autonomously create and refine Python-based tools for visual reasoning, achieving significant performance improvements across benchmarks.

LLMs are increasingly deployed as agents, systems capable of planning,
reasoning, and dynamically calling external tools. However, in visual
reasoning, prior approaches largely remain limited by predefined workflows and
static toolsets. In this report, we present PyVision, an interactive,
multi-turn framework that enables MLLMs to autonomously generate, execute, and
refine Python-based tools tailored to the task at hand, unlocking flexible and
interpretable problem-solving. We develop a taxonomy of the tools created by
PyVision and analyze their usage across a diverse set of benchmarks.
Quantitatively, PyVision achieves consistent performance gains, boosting
GPT-4.1 by +7.8% on V* and Claude-4.0-Sonnet by +31.1% on VLMsAreBlind-mini.
These results point to a broader shift: dynamic tooling allows models not just
to use tools, but to invent them, advancing toward more agentic visual
reasoning.

Source link

What's Hot

The great AI agent acceleration: Why enterprise adoption is happening faster than anyone predicted

TU Wien Rendering #9 – Hard and Soft Shadows

Grok AI uses Musk’s views for some responses

Paper page – PyVision: Agentic Vision with Dynamic Tooling

Paper page – Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Paper page – OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding

Paper page – LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPS

Homeland Security Targets Chicago’s National Museum of Puerto Rican Arts & Culture

1,600-Year-Old Tomb of Mayan City’s Founding King Discovered in Belize

Centre Pompidou Cancels Caribbean Art Show, Raising Controversy

‘Night at the Museum’ Reboot in the Works

The great AI agent acceleration: Why enterprise adoption is happening faster than anyone predicted

TU Wien Rendering #9 – Hard and Soft Shadows

Grok AI uses Musk’s views for some responses

What's Hot

Paper page – PyVision: Agentic Vision with Dynamic Tooling

Related Posts

Subscribe to Updates