arXiv AI

Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning

By Advanced AI EditorJune 30, 2025No Comments2 Mins Read

[Submitted on 16 May 2025 (v1), last revised 27 Jun 2025 (this version, v2)]

View a PDF of the paper titled REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning, by Pawin Taechoyotin and 1 other authors

View PDF
HTML (experimental)

Abstract:AI-based peer review systems tend to produce shallow and overpraising suggestions compared to human feedback. Here, we evaluate how well a reasoning LLM trained with multi-objective reinforcement learning (REMOR) can overcome these limitations. We start by designing a multi-aspect reward function that aligns with human evaluation of reviews. The aspects are related to the review itself (e.g., criticisms, novelty) and the relationship between the review and the manuscript (i.e., relevance). First, we perform supervised fine-tuning of DeepSeek-R1-Distill-Qwen-7B using LoRA on PeerRT, a new dataset of high-quality top AI conference reviews enriched with reasoning traces. We then apply Group Relative Policy Optimization (GRPO) to train two models: REMOR-H (with the human-aligned reward) and REMOR-U (with a uniform reward). Interestingly, the human-aligned reward penalizes aspects typically associated with strong reviews, leading REMOR-U to produce qualitatively more substantive feedback. Our results show that REMOR-U and REMOR-H achieve more than twice the average rewards of human reviews, non-reasoning state-of-the-art agentic multi-modal AI review systems, and general commercial LLM baselines. We found that while the best AI and human reviews are comparable in quality, REMOR avoids the long tail of low-quality human reviews. We discuss how reasoning is key to achieving these improvements and release the Human-aligned Peer Review Reward (HPRR) function, the Peer Review Reasoning-enriched Traces (PeerRT) dataset, and the REMOR models, which we believe can help spur progress in the area.

Submission history

From: Pawin Taechoyotin [view email]
[v1]
Fri, 16 May 2025 22:00:49 UTC (1,408 KB)
[v2]
Fri, 27 Jun 2025 02:48:27 UTC (1,408 KB)

Previous ArticleAI on IBM Power: the platform built for enterprise transformation

Next Article An Artsy Visit To Providence, Rhode Island—‘The Creative Capital’

Advanced AI Editor

Leave A Reply