Paper Page - RefineX: Learning To Refine Pre-training Data At Scale From Expert-Guided Programs

RefineX is a scalable framework for improving the quality of large language model pre-training data through programmatic editing, yielding better performance than alternative methods across various downstream tasks.

The foundational capabilities of large language models (LLMs) are deeply
influenced by the quality of their pre-training corpora. However, enhancing
data quality at scale remains a significant challenge, primarily due to the
trade-off between refinement effectiveness and processing efficiency. While
rule-based filtering remains the dominant paradigm, it typically operates at
the document level and lacks the granularity needed to refine specific content
within documents. Inspired by emerging work such as ProX, we propose
RefineX, a novel framework for large-scale, surgical refinement of
pre-training data through programmatic editing tasks. RefineX enables efficient
and fine-grained data refinement while reliably preserving the diversity and
naturalness of raw text. The core strength of RefineX lies in distilling
high-quality, expert-guided end-to-end refinement results into minimal
edit-based deletion programs. This high-precision distillation pipeline is used
to train an efficient and reliable refine model that can systematically improve
every instance in the corpus at scale. We evaluate RefineX across from-scratch
pre-training at multiple model scales and find that it consistently outperforms
models trained on raw, filtered, or alternatively refined data across diverse
downstream tasks. On the 750M model, RefineX yields 2.6%-7.2% average gains on
lighteval tasks, and achieves comparable performance using significantly fewer
training tokens. Further analysis shows that RefineX reliably enhances text
quality with both high efficiency and precision, outperforming prior approaches
such as end-to-end generation and Prox-C. These results position RefineX as a
scalable, effective, and reliable solution for optimizing pre-training data in
modern LLM pipelines.

Source link

What's Hot

Levi & Korsinsky Reminds C3.ai Investors of the Pending Class Action Lawsuit With a Lead Plaintiff Deadline of October 21, 2025 – AI

USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning – Takara TLDR

British lawmakers accuse Google DeepMind of ‘breach of trust’ over delayed Gemini 2.5 Pro safety report

Paper page – RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning – Takara TLDR

AWorld: Orchestrating the Training Recipe for Agentic AI – Takara TLDR

TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning – Takara TLDR

Woodmere Art Museum Sues Trump Administration Over Canceled IMLS Grant

Barbara Gladstone’s Chelsea Townhouse in NYC Sells for $13.1 M.

Trump Meets with Smithsonian Leader Amid Threats of Content Review

Australian School Faces Pushback over AI Art Course—and More Art News

Levi & Korsinsky Reminds C3.ai Investors of the Pending Class Action Lawsuit With a Lead Plaintiff Deadline of October 21, 2025 – AI

USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning – Takara TLDR

British lawmakers accuse Google DeepMind of ‘breach of trust’ over delayed Gemini 2.5 Pro safety report

What's Hot

Paper page – RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

Related Posts

Subscribe to Updates