Paper Page - DuaShepherd: Integrating Stepwise Correctness And Potential Rewards For Mathematical Reasoning

A novel reward modeling framework DuaShepherd integrates correctness and potential signals into a unified multi-head architecture to enhance LLMs’ mathematical reasoning capabilities and achieve state-of-the-art performance.

In this paper, we propose DuaShepherd, a novel reward modeling framework that
integrates two complementary reward signals, correctness and potential, to
enhance the mathematical reasoning capabilities of Large Language Models
(LLMs). While correctness-based signals emphasize identification of stepwise
errors, potential-based signals focus on the likelihood of reaching the correct
final answer. We developed an automated pipeline for constructing large-scale
reward modeling dataset with both signals. A unified, multi-head architecture
was explored to train the two reward models in a multi-task setup,
demonstrating benefits from learning both correctness and potential in
parallel. By combining these two signals into a compound probability, our model
achieves consistent performance improvements across multiple benchmarks.
Empirical evaluations on MATH500 and ProcessBench confirm that this combined
reward significantly outperforms models trained on either reward type alone,
achieving state-of-the-art performance under comparable resource constraints.

Source link

What's Hot

Watch and Learn: Learning to Use Computers from Online Videos – Takara TLDR

CodeMender from Google DeepMind uses AI to detect bugs and create validated security patches

ChatGPT Now Lets Users Connect With Spotify And Zillow In Chats

Paper page – DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning

Watch and Learn: Learning to Use Computers from Online Videos – Takara TLDR

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights – Takara TLDR

Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails – Takara TLDR

Basquiat Work on Paper Headline’s Phillips’ Frieze Week Sales

Charges Against Isaac Wright ‘to Be Dropped’ After His Arrest by NYPD

What the Los Angeles Wildfires Taught the Art Insurance Industry

Musée d’Orsay Puts Manet on (Mock) Trial for Obscenity

Watch and Learn: Learning to Use Computers from Online Videos – Takara TLDR

CodeMender from Google DeepMind uses AI to detect bugs and create validated security patches

ChatGPT Now Lets Users Connect With Spotify And Zillow In Chats

What's Hot

Paper page – DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning

Related Posts

Subscribe to Updates