Reconstruction Alignment Improves Unified Multimodal Models - Takara TLDR

Unified multimodal models (UMMs) unify visual understanding and generation
within a single architecture. However, conventional training relies on
image-text pairs (or sequences) whose captions are typically sparse and miss
fine-grained visual details–even when they use hundreds of words to describe a
simple image. We introduce Reconstruction Alignment (RecA), a
resource-efficient post-training method that leverages visual understanding
encoder embeddings as dense “text prompts,” providing rich supervision without
captions. Concretely, RecA conditions a UMM on its own visual understanding
embeddings and optimizes it to reconstruct the input image with a
self-supervised reconstruction loss, thereby realigning understanding and
generation. Despite its simplicity, RecA is broadly applicable: across
autoregressive, masked-autoregressive, and diffusion-based UMMs, it
consistently improves generation and editing fidelity. With only 27 GPU-hours,
post-training with RecA substantially improves image generation performance on
GenEval (0.73$\rightarrow$0.90) and DPGBench (80.93$\rightarrow$88.15), while
also boosting editing benchmarks (ImgEdit 3.38$\rightarrow$3.75, GEdit
6.94$\rightarrow$7.25). Notably, RecA surpasses much larger open-source models
and applies broadly across diverse UMM architectures, establishing it as an
efficient and general post-training alignment strategy for UMMs

Source link

What's Hot

Tencent Releases Open Source Image Model HunyuanImage2.1; Aishi Technology Secures $60 Million Funding; Freepik Launches Doubao Seedream 4.0 Image Model_image_the_and

Tencent Hunyuan Image 2.1 Open Source, Advancements in Video Generation and 3D Digital Human Technology_areas_image_the

MLPerf Inference v5.1 Results Land With New Benchmarks and Record Participation

Reconstruction Alignment Improves Unified Multimodal Models – Takara TLDR

Visual Representation Alignment for Multimodal Large Language Models – Takara TLDR

Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding – Takara TLDR

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward – Takara TLDR

Ralph Rugoff to Leave London’s Hayward Gallery After 20 Years

Growing Support for Parthenon Marbles’ Return to Greece, More Art News

Leon Black and Leslie Wexner’s Letters to Jeffrey Epstein Released

School of Visual Arts Transfers Ownership to Nonprofit Alumni Society

Tencent Releases Open Source Image Model HunyuanImage2.1; Aishi Technology Secures $60 Million Funding; Freepik Launches Doubao Seedream 4.0 Image Model_image_the_and

Tencent Hunyuan Image 2.1 Open Source, Advancements in Video Generation and 3D Digital Human Technology_areas_image_the

MLPerf Inference v5.1 Results Land With New Benchmarks and Record Participation

What's Hot

Reconstruction Alignment Improves Unified Multimodal Models – Takara TLDR

Related Posts

Subscribe to Updates