CORTEXA
← Browse
arxivcs.CV2026-07-16

Trajectory-aware Cross-view Geo-localization with Sequential Observations

Tianyi Gao, Jiayu Lin, Danielle Beaulieu, Nathan Jacobs

Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet they overlook a complementary sequential modality: route descriptions -- which capture the same trajectory at a higher level of abstraction and are often the only input available (e.g., a user directing an autonomous vehicle to a pickup point). To bridge this gap, we introduce SeqGeo-VL, a dataset of $\sim$39K video-text-satellite triplets, and TrajLoc, a unified framework capable of processing both video clips and route descriptions. By leveraging both dense visual and abstract linguistic semantics, TrajLoc enables these modalities to mutually reinforce cross-view matching. We further propose TrajMod, a lightweight module that conditions query embeddings on trajectory geometry, yielding spatially-aware representations. Experiments show that TrajLoc achieves substantial gains over state-of-the-art methods on both video and text geo-localization. The project page is available at https://humblegamer.github.io/trajloc/.

View free PDFSource page

Related papers

arxivcs.CV2026-06-30

PiLoT v2: Pixel-to-Orthogonal Map Alignment for Free-view UAV Geo-localization

Xinyi Liu, Xiaoya Cheng, Rouwan Wu, Zhaochen Wang, Shen Yan, Maojun Zhang, et al.

Real-time, drift-free UAV geo-localization is essential for autonomous missions in GNSS-denied environments. The pioneering system, PiLoT, achieves high precision via Neural Pixel-to-3D Registration, aligning UAV video streams with a single rendered reference view from 3D meshes.…

View free PDFSource page
arxivcs.CV2026-06-29

Learning Cross-view Correspondences for Geo-localization on Planetary Surfaces

Hong Minh Nguyen, Marcus Märtens, Tat-Jun Chin

Maintaining global position awareness is a fundamental challenge for planetary surface exploration, since satellite-based positioning systems are unavailable and onboard odometry drifts over time. Although orbital mapping products, such as overhead imagery and terrain-derived map…

View free PDFSource page
arxivcs.CV2026-07-14

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

Mao Chen, Xiangkai Zhang, Zhiyong Liu, Chuankai Liu, Xu Yang

Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it req…

View free PDFSource page
arxivcs.CV2026-07-05

Framework and Multi-modal Dataset for Roadwork Zone Detection and Geo-localization

Zhiran Yan, Yutong Xin, S Shyam Shenoi, Rui Song, Gordon Elger

Autonomous vehicles often rely on high-definition (HD) maps for navigation; however, these maps are not frequently updated and often lack semi-static information, such as temporary roadwork zones, which can significantly alter the road network. This limitation underscores the urg…

View free PDFSource page
arxivcs.CVcs.AI2026-06-29

Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu

Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged reference image (e.g., satellite). Existing approaches heavily rely on 2D appearance matching and are constrained by limited datasets lacking ge…

View free PDFSource page
arxivcs.CV2026-06-29

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning

Xin Zou, Haolin Deng, Yibo Yan, Shuliang Liu, Kening Zheng, Zhiwei Jin, et al.

Inductive biases steer learning toward generalizable solutions by encoding task structure. In this work, we identify a crucial missing bias in MLLMs: cross-view consistency, \textit{i.e.}, semantically invariant views of the same instance should lead to the same answer. Standard…

View free PDFSource page