Paper Page - SkyReels-Audio: Omni Audio-Conditioned Talking Portraits In Video Diffusion Transformers

SkyReels-Audio is a unified framework using pretrained video diffusion transformers for generating high-fidelity and coherent audio-conditioned talking portrait videos, supported by a hybrid curriculum learning strategy and advanced loss mechanisms.

The generation and editing of audio-conditioned talking portraits guided by
multimodal inputs, including text, images, and videos, remains under explored.
In this paper, we present SkyReels-Audio, a unified framework for synthesizing
high-fidelity and temporally coherent talking portrait videos. Built upon
pretrained video diffusion transformers, our framework supports infinite-length
generation and editing, while enabling diverse and controllable conditioning
through multimodal inputs. We employ a hybrid curriculum learning strategy to
progressively align audio with facial motion, enabling fine-grained multimodal
control over long video sequences. To enhance local facial coherence, we
introduce a facial mask loss and an audio-guided classifier-free guidance
mechanism. A sliding-window denoising approach further fuses latent
representations across temporal segments, ensuring visual fidelity and temporal
consistency across extended durations and diverse identities. More importantly,
we construct a dedicated data pipeline for curating high-quality triplets
consisting of synchronized audio, video, and textual descriptions.
Comprehensive benchmark evaluations show that SkyReels-Audio achieves superior
performance in lip-sync accuracy, identity consistency, and realistic facial
dynamics, particularly under complex and challenging conditions.

Source link

What's Hot

Moveworks releases its next-generation copilot, taking action across all business systems using natural language

ASML Invests $1.5 Billion in Mistral AI, Taking Lead Stake in Europe’s Top AI Startup

Meet Blueshoe, the YC-Backed Legal Research Challenger – Artificial Lawyer

Paper page – SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward – Takara TLDR

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions – Takara TLDR

Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling – Takara TLDR

Leon Black and Leslie Wexner’s Letters to Jeffrey Epstein Released

School of Visual Arts Transfers Ownership to Nonprofit Alumni Society

Cristin Tierney Moves Gallery to Tribeca for 15th Anniversary Exhibition

Anne Imhof Reimagines Football Jerseys with Nike

Moveworks releases its next-generation copilot, taking action across all business systems using natural language

ASML Invests $1.5 Billion in Mistral AI, Taking Lead Stake in Europe’s Top AI Startup

Meet Blueshoe, the YC-Backed Legal Research Challenger – Artificial Lawyer

What's Hot

Paper page – SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Related Posts

Subscribe to Updates