UniEval: Unified Holistic Evaluation For Unified Multimodal Understanding And Generation

arXiv:2505.10483v1 Announce Type: cross
Abstract: The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while minimizing model redundancy. However, there is a lack of a unified evaluation framework for these models, which would enable an elegant, simplified, and overall evaluation. Current models conduct evaluations on multiple task-specific benchmarks, but there are significant limitations, such as the lack of overall results, errors from extra evaluation models, reliance on extensive labeled images, benchmarks that lack diversity, and metrics with limited capacity for instruction-following evaluation. To tackle these challenges, we introduce UniEval, the first evaluation framework designed for unified multimodal models without extra models, images, or annotations. This facilitates a simplified and unified evaluation process. The UniEval framework contains a holistic benchmark, UniBench (supports both unified and visual generation models), along with the corresponding UniScore metric. UniBench includes 81 fine-grained tags contributing to high diversity. Experimental results indicate that UniBench is more challenging than existing benchmarks, and UniScore aligns closely with human evaluations, surpassing current metrics. Moreover, we extensively evaluated SoTA unified and visual generation models, uncovering new insights into Univeral’s unique values.

Source link

What's Hot

Sales Plunge 19%! Mercedes Faces Hard Truth and Partners with ‘Doubao’, Can It Turn Things Around This Time?_market_the_’Doubao’

AI Integration Lags Behind the Hype – Artificial Lawyer

Twilio, Palantir Technologies, C3.ai, ZoomInfo, and AppLovin Shares Plummet, What You Need To Know

UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

LTLCrit: A Temporal Logic-based LLM Critic for Safe and Efficient Embodied Agents

From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection

VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots

Smithsonian Closes Museums Amid Government Shutdown

The Rubin Names 2025 Art Prize, Research and Art Projects Grants

Kochi-Muziris Biennial Announces 66 Artists for December Exhibition

Instagram Launches ‘Rings’ Awards for Creators—With KAWS as a Judge

Sales Plunge 19%! Mercedes Faces Hard Truth and Partners with ‘Doubao’, Can It Turn Things Around This Time?_market_the_’Doubao’

AI Integration Lags Behind the Hype – Artificial Lawyer

Twilio, Palantir Technologies, C3.ai, ZoomInfo, and AppLovin Shares Plummet, What You Need To Know

What's Hot

UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

Related Posts

Subscribe to Updates