C1 Advanced Real News

Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training. We propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target. Specifically, we define an example's influence by how much its gradient update reduces the squared distance to the final parameters of a given pretraining run, and estimate this quantity from intermediate checkpoints without retraining. Applying the method to 18 configurations from the Pythia and PolyPythia suites, we find systematic temporal changes in influential data. Early in training, literature-related data are more strongly aligned with the trajectory toward the final parameters, whereas STEM data become more strongly aligned in later stages. This qualitative crossover is broadly consistent across model configurations. Our results provide a tractable trajectory-level view of how influential data change throughout pretraining, complementing influence analyses defined with respect to specific downstream tasks or validation sets.

Cover image for Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
Image: Daily English Reader / Local generated SVG (Project-owned local asset)
5 min read C1

C1 reading

Select any word for its Thai meaning and pronunciation.

0:00 0:00
Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training. We propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target. Specifically, we define an example's influence by how much its gradient update reduces the squared distance to the final parameters of a given pretraining run, and estimate this quantity from intermediate checkpoints without retraining. Applying the method to 18 configurations from the Pythia and PolyPythia suites, we find systematic temporal changes in influential data. Early in training, literature-related data are more strongly aligned with the trajectory toward the final parameters, whereas STEM data become more strongly aligned in later stages. This qualitative crossover is broadly consistent across model configurations. Our results provide a tractable trajectory-level view of how influential data change throughout pretraining, complementing influence analyses defined with respect to specific downstream tasks or validation sets.

ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้

Useful phrases from this story

Measuring training data influence consistentlyCollocation

การวัดผลกระทบของข้อมูลการฝึกอบรมอย่างต่อเนื่อง.

From the storyMeasuring training data influence consistently across language model pretraining is challenging.

pretraining is challengingCollocation

การฝึกซ้อมก่อน เป็นการท้าทาย.

From the storyMeasuring training data influence consistently across language model pretraining is challenging.

is difficult to select downstreamCollocation

ยากที่จะเลือกลงstream.

From the storyIt is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training.

training data influence that doesCollocation

การฝึกอบรมข้อมูลผลกระทบ.

From the storyWe propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target.

selecting a downstream task orCollocation

เลือกงานล่างstream หรือ.

From the storyWe propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target.

Save & Review

Only words saved from this story appear here.