ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: https://github.com/avaxiao/ReToken
C1 reading
Select any word for its Thai meaning and pronunciation.
แปลไทยทั้งบท
สถานการณ์ทางการมองเห็นยาวนานเป็นปัญหาสำหรับรูปแบบของภาษาการมองเห็น: ผลงานลดลงเมื่อจํานวนเครื่องเบี่ยงเบนเพิ่มขึ้น และการแปรรูปของท็อกน์ทั้งหมดพร้อมกันเป็นการคํานวณที่ไม่ทำได้ทำได้ภายใต้ข้อจําจํากัดของจีพีู. เรานําเสนอ ReToken ซึ่งเป็นการพิมพ์ที่ทำได้เรียนรู้ได้แบบเดียว ที่ได้รับการฝึกอบรมเป็นเป้าหมายการค้นหาอย่างชัดเจน ซึ่งเลือกชุดของท็อกน์ภาพที่มีความเกี่ยวข้องกับคําสอบถาม. การฝึกอบรมเพียงบนชุดข้อมูลภาพ-QA ภาพเล็ก ๆ น้อย ๆเท่านั้น, ReToken ผลิตผลประโยชน์ต่อเนื่องผ่านภาพและวิดีโอ benchmarks: ใน Visual Haystacks มันปรับปรุง Qwen3VL-8B โดย 13.4 แต้มและ InternVL3.5 โดย 12.4 แต้ม (>20% ญาติ), และใน LVBench มันโอนศูนย์ถ่ายไปยังวิดีโอยาวสำหรับการเพิ่ม 8.0 แต้ม. กับ Qwen3VL-8B.
ขอบคุณการออกแบบเบาๆ ทั้งการฝึกอบรมและการสรุปวิดีโอยาวๆ เหมาะกับ H100 เดียว. เคล็ดลับทำได้หาได้ที่:.
ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้
Useful phrases from this story
ปรับปรุงท็อคทั้งหมดพร้อมกัน.
From the storyLong visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints.
ไม่สามารถทําการคํานวณได้ภายใต้ GPU.
From the storyLong visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints.
การติดตั้งที่ได้รับการฝึกอบรมอย่างชัดเจน.
From the storyWe present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache.
เติมภาพ KV คาช.
From the storyWe present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache.
มีการฝึกอบรมเพียงเล็กน้อย.
From the storyTrained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B.
Save & Review
Only words saved from this story appear here.