Paper Page - MMHU: A Massive-Scale Multimodal Benchmark For Human Behavior Understanding

A large-scale benchmark, MMHU, is proposed for human behavior analysis in autonomous driving, featuring rich annotations and diverse data sources, and benchmarking multiple tasks including motion prediction and behavior question answering.

Humans are integral components of the transportation ecosystem, and
understanding their behaviors is crucial to facilitating the development of
safe driving systems. Although recent progress has explored various aspects of
human behaviorx2014such as motion, trajectories, and
intentionx2014a comprehensive benchmark for evaluating human
behavior understanding in autonomous driving remains unavailable. In this work,
we propose MMHU, a large-scale benchmark for human behavior analysis
featuring rich annotations, such as human motion and trajectories, text
description for human motions, human intention, and critical behavior labels
relevant to driving safety. Our dataset encompasses 57k human motion clips and
1.73M frames gathered from diverse sources, including established driving
datasets such as Waymo, in-the-wild videos from YouTube, and self-collected
data. A human-in-the-loop annotation pipeline is developed to generate rich
behavior captions. We provide a thorough dataset analysis and benchmark
multiple tasksx2014ranging from motion prediction to motion
generation and human behavior question answeringx2014thereby
offering a broad evaluation suite. Project page :
https://MMHU-Benchmark.github.io.

Source link

What's Hot

The Missing Link in OpenAI’s Deal With Nvidia: Access to Power

Mercor CEO explains how AI affects who gets hired next

EpiCache: Episodic KV Cache Management for Long Conversational Question Answering – Takara TLDR

Paper page – MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding

EpiCache: Episodic KV Cache Management for Long Conversational Question Answering – Takara TLDR

GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning – Takara TLDR

LIMI: Less is More for Agency – Takara TLDR

Court Rules ‘Gender Ideology’ Ban on Art Endowments Unconstitutional

Rural Danish Art Museum Acquires Painting By Artemisia Gentileschi

Dan Nadel Is Expanding American Art History, One Outlier at a Time

Bernard Arnault Says French Wealth Tax Will ‘Destroy’ the Economy

The Missing Link in OpenAI’s Deal With Nvidia: Access to Power

Mercor CEO explains how AI affects who gets hired next

EpiCache: Episodic KV Cache Management for Long Conversational Question Answering – Takara TLDR

What's Hot

Paper page – MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding

Related Posts

Subscribe to Updates