Boosting LLM Exploration via Weak-Model Guidance in RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric perturbation is also neglected. In this work, we propose a simple yet effective approach to preserve the generative diversity of LLMs during RLVR. Instead of relying solely on internal exploration, we force the target model to generate answers based on partial reasoning trajectories generated by a smaller, weaker language models. These unfamiliar prefixes effectively disrupt over-confidence and encourage the exploration of distinct reasoning paths. We empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training. Experiments across multiple mathematical benchmarks show that our method consistently outperforms vanilla RLVR. Notably, the performance gain becomes increasingly pronounced as $k$ scales up, demonstrating a substantial expansion of reasoning coverage. Furthermore, our approach efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.
C1 reading
Select any word for its Thai meaning and pronunciation.
แปลไทยทั้งบท
การเรียนรู้เสริมด้วยรางวัลที่ทำได้ตรวจสอบได้ (RLVR) ช่วยปรับปรุงเหตุผล LLM ได้อย่างสําคัญ แต่มักจะทำให้การประกอบนโยบายลดลง ซึ่งนําไปสู่การกอบคลุมเหตุผลที่คบลง และลดลง pass@$k$ สำหรับ $k$ ใหญ่. ขณะที่วิธีการที่มีอยู่จะลดความล้มเหลวของ entropy นี้ผ่านการปกติแบบอัลโกลิทมิค การปรับปรุงแบบที่ไม่มีปารามเมตรต่าง ๆ ก็ยังถูกละเลย. ในงานนี้ เราเสนอแนวทางที่ง่าย แต่มีประสิทธิภาพ เพื่ออนุรักษ์ความหลากหลายของ LLM ในช่วง RLVR.
แทนที่จะพึ่งพากันเพียงแค่การสํารวจภายใน เราบังคับให้รูปแบบเป้าหมาย สร้างคําตอบ. ปริมาตรที่ไม่คุ้นเคยเหล่านี้ ทำให้เกิดการขัดแย้งความเชื่อมั่นเกิน และส่งเสริมการสํารวจแนวทางการคิดที่แตกต่างกัน. เราพิจารณาการศึกษาความเป็นไปได้ของตัวอักษรก่อนหน้าภายนอก โดยเปิดเผยกลไกของการผลกระทบของความแตกต่างทางการกระจายต่อไดนามิกการสํารวจในการฝึก RLVR.
การทดลองผ่านมาหลายหลักฐานทางคณิตศาสตร์แสดงให้เห็นว่าวิธีของเรามีผลงานที่ดีกว่า RLVR ของวานิลล่า. โดยเฉพาะอย่างยิ่ง การเพิ่มผลงานจะชัดเจนมากขึ้นเมื่อ $k$ เพิ่มขึ้น, แสดงการขยายตัวที่สําคัญของความคุ้มครองเหตุผล. นอกจากนี้ แนวทางของเราช่วยลดความล้มเหลวของ entropy ได้อย่างมีประสิทธิภาพ โดยไม่จําเป็นต้องใช้ SFT เพิ่มเติม การออกแบบรางวัลที่ซับซ้อน หรือการสร้างสาเหตุที่ซับซ้อน.
ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้
Useful phrases from this story
การ เรียน ด้วย ผลตอบแทน ที่ สามารถ ตรวจสอบ ได้.
From the storyReinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$.
เหตุผล แต่มักจะทําให้.
From the storyReinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$.
ส่งผลให้การพิจารณาเหตุผลบําบัด.
From the storyReinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$.
วิธีการที่มีอยู่ลดการเอนตรอปี่นี้.
From the storyWhile existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric perturbation is also neglected.
ยังถูกมองไม่เห็น.
From the storyWhile existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric perturbation is also neglected.
Save & Review
Only words saved from this story appear here.