Paper Page - Whole-Body Conditioned Egocentric Video Prediction

A model trained on real-world egocentric video and body pose predicts video from human actions using an auto-regressive conditional diffusion transformer, evaluated with a hierarchical protocol of tasks.

We train models to Predict Ego-centric Video from human Actions (PEVA), given
the past video and an action represented by the relative 3D body pose. By
conditioning on kinematic pose trajectories, structured by the joint hierarchy
of the body, our model learns to simulate how physical human actions shape the
environment from a first-person point of view. We train an auto-regressive
conditional diffusion transformer on Nymeria, a large-scale dataset of
real-world egocentric video and body pose capture. We further design a
hierarchical evaluation protocol with increasingly challenging tasks, enabling
a comprehensive analysis of the model’s embodied prediction and control
abilities. Our work represents an initial attempt to tackle the challenges of
modeling complex real-world environments and embodied agent behaviors with
video prediction from the perspective of a human.

Source link

What's Hot

MIT rejects Trump compact, first to stand up to partisan demands

Ready or not, enterprises are betting on AI

[Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant)

Paper page – Whole-Body Conditioned Egocentric Video Prediction

Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window – Takara TLDR

First Try Matters: Revisiting the Role of Reflection in Reasoning Models – Takara TLDR

UniVideo: Unified Understanding, Generation, and Editing for Videos – Takara TLDR

The Rubin Names 2025 Art Prize, Research and Art Projects Grants

Kochi-Muziris Biennial Announces 66 Artists for December Exhibition

Instagram Launches ‘Rings’ Awards for Creators—With KAWS as a Judge

Museums Prepare to Close Their Doors as Government Shutdown Continues

MIT rejects Trump compact, first to stand up to partisan demands

Ready or not, enterprises are betting on AI

[Paper Analysis] On the Theoretical Limitations of Embedding-Based Retrieval (Warning: Rant)

What's Hot

Paper page – Whole-Body Conditioned Egocentric Video Prediction

Related Posts

Subscribe to Updates