Towards A Unified View Of Large Language Model Post-Training - Takara TLDR

Two major sources of training data exist for post-training modern language
models: online (model-generated rollouts) data, and offline (human or
other-model demonstrations) data. These two types of data are typically used by
approaches like Reinforcement Learning (RL) and Supervised Fine-Tuning (SFT),
respectively. In this paper, we show that these approaches are not in
contradiction, but are instances of a single optimization process. We derive a
Unified Policy Gradient Estimator, and present the calculations of a wide
spectrum of post-training approaches as the gradient of a common objective
under different data distribution assumptions and various bias-variance
tradeoffs. The gradient estimator is constructed with four interchangeable
parts: stabilization mask, reference policy denominator, advantage estimate,
and likelihood gradient. Motivated by our theoretical findings, we propose
Hybrid Post-Training (HPT), an algorithm that dynamically selects different
training signals. HPT is designed to yield both effective exploitation of
demonstration and stable exploration without sacrificing learned reasoning
patterns. We provide extensive experiments and ablation studies to verify the
effectiveness of our unified theoretical framework and HPT. Across six
mathematical reasoning benchmarks and two out-of-distribution suites, HPT
consistently surpasses strong baselines across models of varying scales and
families.

Source link

What's Hot

AI Futures Project: 2027 AI Forecast Report_The_years_OpenAI

IBM’s Strategic Focus in the Chinese Market Has Shifted, Multinational Tech Giants Favor AI Manufacturing_Dekkers_Yicai

Roblox announces short-form video feed for gameplay clips, new AI tools for creators, and more

Towards a Unified View of Large Language Model Post-Training – Takara TLDR

Delta Activations: A Representation for Finetuned Large Language Models – Takara TLDR

DeepResearch Arena: The First Exam of LLMs’ Research Abilities via Seminar-Grounded Tasks – Takara TLDR

From Editor to Dense Geometry Estimator – Takara TLDR

Basquiats Linked to 1MDB Scandal Auctioned by US Government

US Ambassador to UK Fills Residence with Impressionist Masters

New Code of Ethics Implores UK Museums to End Fossil Fuel Sponsorships

Art Basel Paris Director Clément Delépine to Lead Lafayette Anticipations

AI Futures Project: 2027 AI Forecast Report_The_years_OpenAI

IBM’s Strategic Focus in the Chinese Market Has Shifted, Multinational Tech Giants Favor AI Manufacturing_Dekkers_Yicai

Roblox announces short-form video feed for gameplay clips, new AI tools for creators, and more

What's Hot

Towards a Unified View of Large Language Model Post-Training – Takara TLDR

Related Posts

Subscribe to Updates