Paper Page - FairyGen: Storied Cartoon Video From A Single Child-Drawn Character

FairyGen generates story-driven cartoon videos from a single drawing by disentangling character modeling and background styling, employing MLLM for storyboards, style propagation for consistency, and MMDiT-based diffusion models for motion.

We propose FairyGen, an automatic system for generating story-driven cartoon
videos from a single child’s drawing, while faithfully preserving its unique
artistic style. Unlike previous storytelling methods that primarily focus on
character consistency and basic motion, FairyGen explicitly disentangles
character modeling from stylized background generation and incorporates
cinematic shot design to support expressive and coherent storytelling. Given a
single character sketch, we first employ an MLLM to generate a structured
storyboard with shot-level descriptions that specify environment settings,
character actions, and camera perspectives. To ensure visual consistency, we
introduce a style propagation adapter that captures the character’s visual
style and applies it to the background, faithfully retaining the character’s
full visual identity while synthesizing style-consistent scenes. A shot design
module further enhances visual diversity and cinematic quality through frame
cropping and multi-view synthesis based on the storyboard. To animate the
story, we reconstruct a 3D proxy of the character to derive physically
plausible motion sequences, which are then used to fine-tune an MMDiT-based
image-to-video diffusion model. We further propose a two-stage motion
customization adapter: the first stage learns appearance features from
temporally unordered frames, disentangling identity from motion; the second
stage models temporal dynamics using a timestep-shift strategy with frozen
identity weights. Once trained, FairyGen directly renders diverse and coherent
video scenes aligned with the storyboard. Extensive experiments demonstrate
that our system produces animations that are stylistically faithful,
narratively structured natural motion, highlighting its potential for
personalized and engaging story animation. The code will be available at
https://github.com/GVCLab/FairyGen

Source link

What's Hot

Reinforcing Diffusion Models by Direct Group Preference Optimization – Takara TLDR

it takes more than chips to win the AI race

Alibaba’s Artificial Intelligence (AI) Push: Could This Be China’s Best Answer to Nvidia?

Paper page – FairyGen: Storied Cartoon Video from a Single Child-Drawn Character

Reinforcing Diffusion Models by Direct Group Preference Optimization – Takara TLDR

Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency – Takara TLDR

DeepPrune: Parallel Scaling without Inter-trace Redundancy – Takara TLDR

The Rubin Names 2025 Art Prize, Research and Art Projects Grants

Kochi-Muziris Biennial Announces 66 Artists for December Exhibition

Frieze to Launch Abu Dhabi Fair in November 2026

Jeff Koons Returns to Gagosian with First New York Show in Seven Years

Reinforcing Diffusion Models by Direct Group Preference Optimization – Takara TLDR

it takes more than chips to win the AI race

Alibaba’s Artificial Intelligence (AI) Push: Could This Be China’s Best Answer to Nvidia?

What's Hot

Paper page – FairyGen: Storied Cartoon Video from a Single Child-Drawn Character

Related Posts

Subscribe to Updates