CORTEXA
← Browse
arxivcs.CVcs.AI2026-07-17

Geometry-Enhanced Portion Estimation for Multimodal LLMs

Lin Liao, Peng Li

Image-based dietary assessment promises to replace costly, bias-prone manual recalls, but portion estimation remains a major blocker. Multimodal LLMs (MLLMs) recognize a wide range of foods zero-shot in uncontrolled photos, yet they are weak at portion estimation -- a gap we measure across the current frontier (Gemini, GPT, and Claude flagships alike). We present a method that enhances a frozen, commercial MLLM with an accurate portion head: a small geometry-enhanced network on a frozen DINOv2 backbone with a structured softmax-ownership volume, consuming the MLLM's per-food name, bounding box, and density range -- no depth sensor, no MLLM fine-tuning. Evaluated fully open-vocabulary on three real-world benchmarks, the head cuts per-food portion error by 33-41% relative to the MLLM alone, outperforms every flagship MLLM's direct estimates, and surpasses each benchmark's originally published image-only model at its own reported metric.

View free PDFSource page

Related papers

arxivcs.CVcs.AIcs.CLcs.MM2026-07-15

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs

Zhixiao Zheng, Zheren Fu, Zhiyuan Yao, Chunxiao Liu, Dongming Zhang, Zhendong Mao

Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual hallucinations, content fabrication, and unfaithful reasoning, which substantially undermine their faithfulness and practical utility. Alignment me…

View free PDFSource page
arxivcs.CVcs.AI2026-07-15

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs

Taorui Wang, Wei Xia, Hui Ma, Zijia Song, Jiayu Zhang, Zeheng Wang, et al.

Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise. While Multimodal Large Language Models (MLLMs) excel at general video understanding, they inherently str…

View free PDFSource page
arxivcs.CVcs.AI2026-07-02

K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

Khush Attarde, Yusuf Ali, Megha Thukral, Divye Bhutani, Thomas Ploetz, Zsolt Kira

MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, application of these models lies in understanding and modeling animal-centric scenarios. As animals are integral to millions of househ…

View free PDFSource page
arxivcs.CVcs.AI2026-07-02

GAP-GDRNet: Geometry-aware monocular 6D pose estimation for spacecraft using synthetic geometric supervision

Zongwu Xie, Yonglong Zhang, Yifan Yang, Yang Liu, Guanghu Xie

Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion. This article presents GAP-GDRNet, a geometry-aware RGB framework built on GDR-Net for a single-target synthetic spacecraft benchmark. The method s…

View free PDFSource page
arxivcs.CVcs.AIcs.CY2026-07-03

Responsibility Distribution Estimation in Ego-View Accident Videos with Multimodal Large Language Models

Ryosei Tamura, Andrew Shin

Recent studies on multimodal traffic accident understanding have mainly relied on infrastructure-camera footage, satellite imagery, or structured crash records. However, such data sources are costly to deploy and maintain at large scale, and they cannot objectively capture what t…

View free PDFSource page
arxivcs.CVcs.AI2026-07-08

Beyond Thermal Imaging: Inferring Thermophysical Properties from Time-Resolved Thermal Observations

Chenghao Xu, Malcolm Mielle, Olga Fink

Inferring latent physical properties from sensory observations is a fundamental challenge in machine perception. Among available sensing modalities, thermal imaging is particularly promising because temperature evolution is directly governed by heat-transfer physics and therefore…

View free PDFSource page