CORTEXA
← Browse
arxivcs.CV2026-06-30

CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization

Bo Wu, Ruoshen Mo, Justin Yue, Yanyu Zhang, Janice Nguyen, Guoyuan Wu, Amit Roy-Chowdhury, Matthew J. Barth, Hang Qiu

Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication bandwidth and its dynamics, heterogeneous sensing modalities, and scalability beyond a single cooperative partner. In this paper, we introduce CooperScene, a high-fidelity cooperative autonomy dataset with real-world C-V2X communication characterization. The dataset is organized into diverse scenes, including intersections, highway ramps, and parking lots. These scenes involve three connected and autonomous vehicles (CAVs) and one infrastructure roadside unit (RSU), all equipped with multi-modal sensors and commercial off-the-shelf C-V2X communication radios. All scenes are annotated with globally consistent 3D labels at 10 Hz, totaling 344K objects across 59K frames, underpinned by tight sensor- and agent-synchronization, centimeter-level localization and spatial alignment, precise cross-modality calibration, and 3GPP-standard-compliant C-V2X communication. CooperScene establishes a rigorous benchmark for evaluating multi-agent scaling and actual performance in real-world deployable settings. Project website for data and benchmark: https://cisl.ucr.edu/CooperScene

View free PDFSource page

Related papers

arxivcs.CVcs.AI2026-07-01

GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception

Xiao Zhao, Chang Liu, Mingxu Zhu, Zheyuan Zhang, Linna Song, Qingliang Luo, et al.

The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception. However, the discrete grid representation of BEV leads to significant detail loss and limits featu…

View free PDFSource page
arxivcs.CV2026-07-08

Time Imprint: Learning Time-Aware Representations in Multi-Modal Knowledge Graphs

Pengyu Zhang, Klim Zaporojets, Congfeng Cao, Jia-Hong Huang, Paul Groth

Multi-Modal Knowledge Graphs (MMKGs) enrich entities with multiple modalities such as text and images, yet entities with highly similar multi-modal features remain difficult to distinguish. Temporal information of an entity can serve as an additional modality to disambiguate such…

View free PDFSource page
arxivcs.CV2026-07-16

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

Xiao Wang, Bing Wang, Bin Yang, Cuiqun Chen, Xin Xu, Mang Ye

Person re-identification (ReID) serves as a critical component in intelligent surveillance systems, aiming to match identities across disjoint camera networks. While traditional methods primarily rely on single-modal RGB imagery, they are often constrained by environmental challe…

View free PDFSource page
arxivcs.CV2026-07-08

Two-Stage Multi-Modal Fusion with Adaptive Alignment for Action Quality Assessment

Kanglei Zhou, Ruizhi Cai, Xinning Wang, Yijian Zheng, Liyuan Wang, Jianguo Li, et al.

Action Quality Assessment (AQA) aims to evaluate how well a person performs a movement, which is essential in applications such as sports scoring, skill assessment, and healthcare. However, unimodal approaches often struggle to capture subtle cues of movement quality in real-worl…

View free PDFSource page
arxivcs.CVcs.AIcs.MM2026-07-10

Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

Peipei Zhu, Yueqing Niu, Lin Zhu, Guanchong Niu, Yang Yu, Zheng Li

Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos. To address these limitations, we propose EVAD,…

View free PDFSource page
arxivcs.CV2026-06-30

RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception

Shaozu Ding, Linan Song, Marco De Vincenzi, Dajiang Suo

LiDAR has increasingly been integrated into traffic cameras to expand coverage and mitigate occlusion in roadside cooperative perception. However, how unimodal and camera-LiDAR fusion architectures behave under variations in LiDAR point sparsity induced by sensor configurations a…

View free PDFSource page