BESPOKE: Benchmark For Search-Augmented Large Language Model Personalization Via Diagnostic Feedback - Takara TLDR

Search-augmented large language models (LLMs) have advanced
information-seeking tasks by integrating retrieval into generation, reducing
users’ cognitive burden compared to traditional search systems. Yet they remain
insufficient for fully addressing diverse user needs, which requires
recognizing how the same query can reflect different intents across users and
delivering information in preferred forms. While recent systems such as ChatGPT
and Gemini attempt personalization by leveraging user histories, systematic
evaluation of such personalization is under-explored. To address this gap, we
propose BESPOKE, the realistic benchmark for evaluating personalization in
search-augmented LLMs. BESPOKE is designed to be both realistic, by collecting
authentic chat and search histories directly from humans, and diagnostic, by
pairing responses with fine-grained preference scores and feedback. The
benchmark is constructed through long-term, deeply engaged human annotation,
where human annotators contributed their own histories, authored queries with
detailed information needs, and evaluated responses with scores and diagnostic
feedback. Leveraging BESPOKE, we conduct systematic analyses that reveal key
requirements for effective personalization in information-seeking tasks,
providing a foundation for fine-grained evaluation of personalized
search-augmented LLMs. Our code and data are available at
https://augustinlib.github.io/BESPOKE/.

Source link

What's Hot

how Huawei and DeepSeek are helping China break reliance on US chips

How South Korea plans to best OpenAI, Google, others with homegrown AI

Tesla pleads with Trump White House not to bail on crucial climate standards

BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback – Takara TLDR

Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution – Takara TLDR

MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning – Takara TLDR

CHARM: Control-point-based 3D Anime Hairstyle Auto-Regressive Modeling – Takara TLDR

Judge Rejects Ronald Perelman’s $400 M. Art Insurance Claim

Drag Queen Alexis Stone Became the Mona Lisa for Milan Fashion Show

Steve McQueen’s Granddaughter Lawsuit for $68 M. Pollock Painting

Lisa Phillips, Longtime Director of New York’s New Museum, to Retire

how Huawei and DeepSeek are helping China break reliance on US chips

How South Korea plans to best OpenAI, Google, others with homegrown AI

Tesla pleads with Trump White House not to bail on crucial climate standards

What's Hot

BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback – Takara TLDR

Related Posts

Subscribe to Updates