Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining
Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training. We propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target. Specifically, we define an example's influence by how much its gradient update reduces the squared distance to the final parameters of a given pretraining run, and estimate this quantity from intermediate checkpoints without retraining. Applying the method to 18 configurations from the Pythia and PolyPythia suites, we find systematic temporal changes in influential data. Early in training, literature-related data are more strongly aligned with the trajectory toward the final parameters, whereas STEM data become more strongly aligned in later stages. This qualitative crossover is broadly consistent across model configurations. Our results provide a tractable trajectory-level view of how influential data change throughout pretraining, complementing influence analyses defined with respect to specific downstream tasks or validation sets.
C1 reading
Select any word for its Thai meaning and pronunciation.
แปลไทยทั้งบท
การวัดผลกระทบของข้อมูลการฝึกอบรมอย่างสม่ําเสมอ ผ่านรูปแบบการฝึกอบรมก่อนการฝึกอบรม เป็นความท้าทาย. มันยากที่จะเลือกงานล่วงหน้าหรือชุดการยืนยันที่แสดงถึงความทำได้ทั่วไปของตัวอย่าง และการพึ่งพาการทำงานในจุดตรวจกลางทำให้การเปรียบเทียบระหว่างการฝึกอบรมซับซ้อน. เราเสนอวัดการส่งผลกระทบของข้อมูลการอบรมที่ไม่จําเป็นต้องเลือกงานล่วงหน้าหรือการยืนยันที่ตั้งเป็นเป้าหมายการมอบหมาย.
โดยเฉพาะอย่างยิ่ง เรากําหนดผลกระทบของตัวอย่างโดยการปรับปรุง gradient ของมันลดระยะสี่ถึงปาร์เมตรสุดท้ายของการแข่งขันก่อนการฝึกอบรม และประเมินจํานวนนี้จากจุดตรวจกลางโดยไม่ต้องฝึกซ้อมใหม่. โดยนําวิธีนี้ไปใช้กับ 18 การประกอบจาก Suite Pythia และ PolyPythia เราพบกับการเปลี่ยนแปลงระยะเวลาในข้อมูลที่มีผลกระทบ. ในช่วงต้นของการฝึกอบรม ข้อมูลที่เกี่ยวข้องกับวรรณคดี ถูกสอดคล้องกันอย่างแข็งแกร่งกับเส้นทางไปสู่ปารามิเตอร์สุดท้าย ขณะที่ ข้อมูล STEM จะถูกสอดคล้องอย่างแข็งแกร่งในช่วงหลัง.
การตัดแยกทางคุณภาพนี้มีความสอดคล้องตามรูปแบบทั่วไป. ผลงานของเราให้การมองเห็นระดับเส้นทางที่ทำได้แก้ไขได้เกี่ยวกับการเปลี่ยนแปลงข้อมูลที่มีผลกระทบในช่วงการอบรมก่อนการอบรม โดยเติมเต็มการวิเคราะห์ผลกระทบที่กําหนดในส่วนของงานล่วงหน้าหรือชุดการยืนยันเฉพาะเจาะจง.
ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้
Useful phrases from this story
การวัดผลกระทบของข้อมูลการฝึกอบรมอย่างต่อเนื่อง.
From the storyMeasuring training data influence consistently across language model pretraining is challenging.
การฝึกซ้อมก่อน เป็นการท้าทาย.
From the storyMeasuring training data influence consistently across language model pretraining is challenging.
ยากที่จะเลือกลงstream.
From the storyIt is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training.
การฝึกอบรมข้อมูลผลกระทบ.
From the storyWe propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target.
เลือกงานล่างstream หรือ.
From the storyWe propose a measure of training data influence that does not require selecting a downstream task or validation set as the attribution target.
Save & Review
Only words saved from this story appear here.