CORTEXA
← Browse
arxivcs.CV2026-07-03

Natural Language Camera Movement Understanding

Yuwen Tan, Joey Huang, Jin Huang, Haoxiang Li, Boqing Gong

Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing vision-language models (VLMs) fail this task in surprising ways, frequently confusing translation with rotation, left with right, and object movement with camera movement. To address these limitations, we establish natural language camera movement understanding as a standalone research task. We introduce a two-level cinematographic taxonomy and an extensive, atomic benchmark featuring both real and synthetic videos. Furthermore, we curate a large-scale, multi-source training set enhanced by targeted camera movement augmentation. Our fine-tuned VLM-8B outperforms Gemini 3.1 Pro by 10% and 11% on our benchmark's real and synthetic videos, respectively. Despite these gains, a significant gap remains relative to human performance, underscoring the need to promote and facilitate future research on natural language camera movement understanding.

View free PDFSource page

Related papers

arxivcs.CV2026-07-01

GenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language Models

Hongkuan Zhou, Tristan Rehm, Nadeem Nazer, Lavdim Halilaj, Jingcheng Wu, Steffen Staab

Industrial inspection requires more than binary anomaly detection: a practical system should determine whether an anomaly exists, localize the defective region, identify the defect type, and provide interpretable visual evidence. Existing CLIP-based methods detect and localize an…

View free PDFSource page
arxivcs.CVcs.AIcs.LG2026-06-28

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking

Xiao Wang, Liye Jin, Dan Xu, Yuehang Li, Lan Chen, Yaowei Wang, et al.

Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially boost tracking accuracy and robustness. Existing studies have verified that adaptively optimizing textual descriptions throughout the tracking…

View free PDFSource page
arxivcs.AIcs.CV2026-07-06

ASSEMCAD: Production-Ready CAD Assembly Generation from Natural Language

Yurui Dong, Shu Zou, Siqi Li, Nianchen Deng, Hongbin Zhou, Xuemeng Yang, et al.

Recent advances in large language models and programmatic CAD have significantly improved Text-to-CAD generation for individual parts. However, production-ready mechanical assembly generation remains largely unsolved. Unlike single-part modeling, assemblies require coordinated re…

View free PDFSource page
arxivcs.CV2026-06-29

3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

Deyin Liu, Jicheng Xu, Lin Yuanbo Wu, Xiaowei Zhao, Xiatian Zhu, Zhe Jin, et al.

Human image animation, which aims to generate a video of a reference subject following a provided action sequence, has received increasing research interest. With the development of diffusion-based/flow-based video foundation models, existing animation works have began to upgrade…

View free PDFSource page
arxivcs.CV2026-07-02

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions

Haoran Wang, Mohit Mendiratta, Christian Theobalt, Adam Kortylewski

Precise control of 3D facial expressions from text is crucial for virtual avatars, animation, and human-computer interaction, yet existing text-to-3D methods jointly generate identity, expression, and texture, making fine-grained expression control difficult. We instead formulate…

View free PDFSource page
arxivcs.CVcs.CL2026-07-18

Can Multimodal Large Language Models Understand OCT?

Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li, Zihan Nie, Kailin Jiang, et al.

Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding…

View free PDFSource page