Paper page - Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Large Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optimization framework to address this, employing a closed-loop system where LLMs iteratively refine code based on empirical performance feedback from an execution sandbox. We explore three
training strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO).

Experiments on our Venus dataset and the APPS benchmark show that SFT and DPO rapidly saturate
in efficiency gains. In contrast, GRPO, using reinforcement learning (RL) with execution feedback, continuously optimizes code performance, significantly boosting both PASS@1 (from 47% to 62%) and the likelihood of outperforming human submissions in efficiency (from 31% to 45%). Our work demonstrates effective test-time code efficiency improvement and critically reveals the power of RL in teaching LLMs to truly self-improve code efficiency.

Source link

What's Hot

Employer Branding Grows Companies | Recruiting News Network

Paper page – HOComp: Interaction-Aware Human-Object Composition

DeepSeek Predicts DOGE, BONK And WIF Prices For End Of 2025

Paper page – Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Paper page – HOComp: Interaction-Aware Human-Object Composition

Paper page – Does More Inference-Time Compute Really Help Robustness?

Paper page – RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback

Barnes Foundation Online Learning Platform Expands to Penn Museum

Archaeologists Identify 5,500-Year-Old Megalithic Tombs in Poland

Phillips to Debut ‘First-of-its Kind’ Priority Bidding Structure

3,800-Year-Old Warrior’s Tomb Unearthed in Azerbaijan

Employer Branding Grows Companies | Recruiting News Network

Paper page – HOComp: Interaction-Aware Human-Object Composition

DeepSeek Predicts DOGE, BONK And WIF Prices For End Of 2025

What's Hot

Paper page – Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

Related Posts

Subscribe to Updates