CORTEXA
← Browse
arxivcs.CV2026-07-16

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

Xiao Lin, Xiaohu Huang, Kai Han

Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, different models provide complementary spatial priors that benefit different tasks. Motivated by this, we propose $\textbf{ViPS}$, a novel multi-model prior framework designed to fully unleash the potential of incorporating multiple $\textbf{Vi}$sual $\textbf{P}$riors from diverse models into MLLMs for $\textbf{S}$patial understanding. Specifically, ViPS introduces an Efficient Prior Proxy to generate multiple foundational priors with minimal inference overhead, and a Dynamic Prior Fusion mechanism to achieve harmonious and context-aware prior fusion and injection from the prior proxies. Extensive experiments demonstrate that ViPS successfully harmonizes diverse visual priors, establishing new state-of-the-art performance across multiple complex spatial reasoning and 3D spatial understanding benchmarks. Project page: https://visual-ai.github.io/vips

View free PDFSource page

Related papers

arxivcs.CV2026-06-30

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

Ziqi Li, Zijian Chen, Tingzhu Chen, Guangtao Zhai

Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form textual coherence and contextual dependencies embedded in complete divination charges. Recently, the powerful visual perception…

View free PDFSource page
arxivcs.CVcs.AI2026-06-28

ReMAP-PET: Beyond Visual Understanding -- Learning Region-Guided Metabolic Alignment Semantics from Brain PET

Dasen Dai, Yanteng Zhang, Shuoqi Li, Yuxiang Wei, Hongjie Yu, Qingxin Zhang, et al.

Positron Emission Tomography (PET) reveals brain metabolism and is clinically central to neurodegenerative disease assessment, yet existing 3D brain foundation models treat PET as generic volumetric data, missing the structured regional metabolic information that distinguishes it…

View free PDFSource page
arxivcs.CVcs.AI2026-07-15

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

Hao Li, Han Fang, Zixin Pan, Xin Wei, Hongbo Sun, Jinglin Xu, et al.

Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing methods primarily rely on symbolic text tokens, which inherently lack the fidelity to represent contin…

View free PDFSource page
arxivcs.CVcs.AI2026-06-29

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

Yuxuan Fan, Gyusik Seo, Jing Hao, Jaemin Cho, Mohit Bansal, Jaehong Yoon

Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing,…

View free PDFSource page
arxivcs.AIcs.CV2026-07-20

Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding

Zhiwei Wei, Yonghe Sun, Zhenjia Liu, Wenjia Xu, Chao He, Weihua Dong, et al.

Spatial understanding is crucial for foundation models (FMs), and maps have long helped humans organize and reason about geographic information. This study examines whether choropleth maps remain useful for machine spatial understanding when models can directly process structured…

View free PDFSource page
arxivcs.CV2026-07-13

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

Kerui Chen, Jinglu Wang, Xiaoyi Zhang, Yan Lu

Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, spo…

View free PDFSource page