Paper Page - Robust Preference Optimization Via Dynamic Target Margins

The paper introduces γ-PO, a dynamic target margin preference optimization algorithm that enhances Large Language Models’ alignment by adjusting reward margins at the pairwise level, leading to improved performance with minimal impact on training.

The alignment of Large Language Models (LLMs) is crucial for ensuring their
safety and reliability in practical applications. Direct Preference
Optimization (DPO) has emerged as an efficient method that directly optimizes
models using preference pairs, significantly reducing resource demands.
However, the effectiveness of DPO heavily depends on the data quality, which is
frequently compromised by noise. In this work, we propose gamma-PO, a
dynamic target margin preference optimization algorithm that adjust reward
margins at the pairwise level. By introducing instance-specific margin
calibration, gamma-PO strategically prioritizes high-confidence pairs (those
demonstrating higher reward margins) while suppressing potential noise from
ambiguous pairs. Moreover, gamma-PO is a plug-and-play method, compatible
with variants of DPO that rely on reward margin between preference pairs.
Across benchmarks such as AlpacaEval2 and Arena-Hard, gamma-PO achieves an
average 4.4\% improvement over other baselines, setting new benchmarks for
state-of-the-art performance. Additionally, gamma-PO requires minimal code
changes and has a negligible impact on training efficiency, making it a robust
solution for enhancing LLMs alignment. Our codes are available at
https://github.com/sunjie279/gammaPO{https://github.com/sunjie279/gammaPO}.

Source link

What's Hot

GIST and MIT Launch Full-Scale Research on Human-Centered Physical AI Interaction

Distyl AI Raises $175M Series B At $1.8B Valuation, Up 9x From Last Funding

The Oakland Ballers let an AI manage the team. What could go wrong?

Paper page – Robust Preference Optimization via Dynamic Target Margins

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue – Takara TLDR

RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes – Takara TLDR

BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent – Takara TLDR

St. Patrick’s Cathedral Unveils Monumental Mural by Adam Cvijanovic

Three Loaned Banksy Works Incite Dispute Between England and Italy

Major Collection of Old Masters Paintings Could Be Fractionalized

100 Must-See Artworks at the Metropolitan Museum of Art

GIST and MIT Launch Full-Scale Research on Human-Centered Physical AI Interaction

Distyl AI Raises $175M Series B At $1.8B Valuation, Up 9x From Last Funding

The Oakland Ballers let an AI manage the team. What could go wrong?

What's Hot

Paper page – Robust Preference Optimization via Dynamic Target Margins

Related Posts

Subscribe to Updates