arxivcs.CVcs.AI2026-07-04
Attending to Multimodal Generation One Token at a Time
Varun Gupta, Vineet Gandhi, Makarand Tapaswi
Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving context. Prior work on interpretability has focused on individual layers and circuits (where), leaving the token-level dynamics of multimodal…