Drawing Conclusions From Draws: Rethinking Preference Semantics In Arena-Style LLM Evaluation - Takara TLDR

In arena-style evaluation of large language models (LLMs), two LLMs respond
to a user query, and the user chooses the winning response or deems the
“battle” a draw, resulting in an adjustment to the ratings of both models. The
prevailing approach for modeling these rating dynamics is to view battles as
two-player game matches, as in chess, and apply the Elo rating system and its
derivatives. In this paper, we critically examine this paradigm. Specifically,
we question whether a draw genuinely means that the two models are equal and
hence whether their ratings should be equalized. Instead, we conjecture that
draws are more indicative of query difficulty: if the query is too easy, then
both models are more likely to succeed equally. On three real-world arena
datasets, we show that ignoring rating updates for draws yields a 1-3% relative
increase in battle outcome prediction accuracy (which includes draws) for all
four rating systems studied. Further analyses suggest that draws occur more for
queries rated as very easy and those as highly objective, with risk ratios of
1.37 and 1.35, respectively. We recommend future rating systems to reconsider
existing draw semantics and to account for query properties in rating updates.

Source link

What's Hot

A Rigorous Benchmark with Multidimensional Evaluation for Deep Research Agents: From Answers to Reports – Takara TLDR

Stocks to Gain From Quantum Computing in 2025: MSFT, IBM, QBTS, IONQ

StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? – Takara TLDR

Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation – Takara TLDR

A Rigorous Benchmark with Multidimensional Evaluation for Deep Research Agents: From Answers to Reports – Takara TLDR

StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? – Takara TLDR

RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning – Takara TLDR

Record Exec and Art Collector Gets Over 4 Years

Chicago’s Art Scene Offers a Beacon of Hope for Artists and Dealers

New Archaeological Research Reveals Life in Pompeii Post-Eruption

Director Fired After Declining to Give Trump Sword for King Charles

A Rigorous Benchmark with Multidimensional Evaluation for Deep Research Agents: From Answers to Reports – Takara TLDR

Stocks to Gain From Quantum Computing in 2025: MSFT, IBM, QBTS, IONQ

StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? – Takara TLDR

What's Hot

Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation – Takara TLDR

Related Posts

Subscribe to Updates