C1 Advanced Real News

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: https://github.com/avaxiao/ReToken

Cover image for ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Image: Daily English Reader / Local generated SVG (Project-owned local asset)
5 min read C1

C1 reading

Select any word for its Thai meaning and pronunciation.

0:00 0:00
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at:

ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้

Useful phrases from this story

processing all tokens at onceCollocation

ปรับปรุงท็อคทั้งหมดพร้อมกัน.

From the storyLong visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints.

is computationally infeasible under GPUCollocation

ไม่สามารถทําการคํานวณได้ภายใต้ GPU.

From the storyLong visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints.

embedding trained as an explicitCollocation

การติดตั้งที่ได้รับการฝึกอบรมอย่างชัดเจน.

From the storyWe present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache.

filled visual KV cacheCollocation

เติมภาพ KV คาช.

From the storyWe present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache.

Trained on only a smallCollocation

มีการฝึกอบรมเพียงเล็กน้อย.

From the storyTrained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B.

Save & Review

Only words saved from this story appear here.