ResearcharXiv AI / CL
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Summary
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints.
Original Article
Captured source content or English translation, normalized into this reading format.
This story does not yet have captured source text. Open the source link to read it.
Region
Global
Heat Score
81
Category
Research
Language
en
