Paper Page - Synthetic Data RL: Task Definition Is All You Need

Synthetic Data RL enhances foundation models through reinforcement learning using only synthetic data, achieving performance comparable to models trained with full human-labeled data.

Reinforcement learning (RL) is a powerful way to adapt foundation models to
specialized tasks, but its reliance on large-scale human-labeled data limits
broad adoption. We introduce Synthetic Data RL, a simple and general framework
that reinforcement fine-tunes models using only synthetic data generated from a
task definition. Our method first generates question and answer pairs from the
task definition and retrieved documents, then adapts the difficulty of the
question based on model solvability, and selects questions using the average
pass rate of the model across samples for RL training. On Qwen-2.5-7B, our
method achieves a 29.2% absolute improvement over the base model on GSM8K (+2.9
pp vs. instruction-tuned, +6.6 pp vs. Self-Instruct), 8.7% on MATH, 13.1% on
GPQA (+7.0 pp vs. SynthLLM), 8.9% on MedQA, 17.7% on CQA (law) and 13.7% on CFA
(finance). It surpasses supervised fine-tuning under the same data budget and
nearly matches RL with full human data across datasets (e.g., +17.2 pp on
GSM8K). Adding 100 human demonstrations improves the performance of GSM8K only
by 0.4 pp, showing a limited added value. By reducing human data annotation,
Synthetic Data RL enables scalable and efficient RL-based model adaptation.
Code and demos are available at https://github.com/gydpku/Data_Synthesis_RL/.

Source link

What's Hot

AI Agents + What’s Next for Legal Judgment – Artificial Lawyer

P3-SAM: Native 3D Part Segmentation – Takara TLDR

Stability AI Launches Stable Audio 2.5 with Enterprise-Grade Speed and Creative Control

Paper page – Synthetic Data RL: Task Definition Is All You Need

P3-SAM: Native 3D Part Segmentation – Takara TLDR

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants – Takara TLDR

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning – Takara TLDR

Christie’s Will Auction The First Calculating Machine In History

The Art Market Isn’t Dying. The Way We Write About It Might Be.

Banksy Mural of Judge Beating Protestor Removed by Courts Service

Death of Matthew Christopher Pietras Ruled a Suicide

AI Agents + What’s Next for Legal Judgment – Artificial Lawyer

P3-SAM: Native 3D Part Segmentation – Takara TLDR

Stability AI Launches Stable Audio 2.5 with Enterprise-Grade Speed and Creative Control

What's Hot

Paper page – Synthetic Data RL: Task Definition Is All You Need

Related Posts

Subscribe to Updates