arXiv CS AI#tech
Thinking with Visual Groundingtranslating…
Factuality: 90/100USACornell University
arXiv:2606.16122v1 Announce Type: new
Abstract: Visual thinking should not only sound right; it should show its evidence. While recent vision-language models (VLMs) can produce natural-language reasoning traces, these traces often leave the supporting image regions implicit, making them hard to verify and difficult to supervise. We introduce visually grounded thinking, a reasoning process in which models interleave natural-language thoughts with explicit point or box groundings of the visual ev