CORTEXA
← Browse
arxivcs.CV2026-07-02

Sparse-Aware Vector Quantization for Bandwidth-Efficient Collaborative 3D Semantic Occupancy Prediction

Feng Li, Chaokun Zhang, Gong Chen

Collaborative perception extends single-agent perception by enabling multiple vehicles to exchange complementary perceptual information. However, it introduces an inherent trade-off between perception gain and communication overhead, which is particularly severe for 3D semantic occupancy prediction that relies on fine-grained spatial structures. Existing methods typically compress 3D features into 2D, causing severe spatial information loss, or transmit dense 3D representations, hindering real-world deployment. To overcome these limitations, we propose a bandwidth-efficient collaborative Vector Quantization Semantic Occupancy Prediction (VQSOP) framework. VQSOP employs a Sparse-Aware Vector Quantization (SAVQ) mechanism that exploits 3D scene sparsity to compactly encode informative regions, drastically reducing communication overhead while preserving complete geometric context. Furthermore, to enhance structural consistency and feature continuity, we design a Dual-Branch Adaptive Spatial Refinement (ASR) module that dynamically fuses local high-frequency details with broad contextual semantics. Extensive experiments demonstrate that our approach achieves state-of-the-art performance while reducing communication volume by up to 82x.

View free PDFSource page

Related papers

arxivcs.CV2026-07-04

FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

Dubing Chen, Huan Zheng, Tianyi Yan, Yucheng Zhou, Runzhou Tao, Zhongying Qiu, et al.

Vision-based 3D occupancy prediction fundamentally relies on the 2D-to-3D view transformation. Current paradigms predominantly utilize explicit physical projection, which artificially restricts the routing matrix to strict, sparse camera rays. While computationally efficient, thi…

View free PDFSource page
arxivcs.CV2026-06-30

Semantic Occupancy Prediction with Dual Range-Voxel Representation

Sitao Chen, Zhuangwei Zhuang, Hui Luo, Lizhao Liu, Qingyao Wu, Mingkui Tan

LiDAR-based 3D semantic occupancy prediction, which aims to provide accurate and comprehensive scene representation, is crucial for autonomous driving systems. As point clouds suffer from sparsity and incompleteness, leading to insufficient semantic learning and difficult occupan…

View free PDFSource page
arxivcs.CV2026-07-06

SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction

Pin Tang, Zhongdao Wang, Guoqing Wang, Xiangxuan Ren, Chao Ma

Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty space, while BEV and TPV projections compromise fine-grained 3D structure. Fully sparse representations offer an attractive altern…

View free PDFSource page
arxivcs.CV2026-07-20

VGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy Prediction

Junhong Lin, Xianda Guo, Kangli Wang, Yuqi Ye, Xiaoyu Liang, Yanlun Peng, et al.

Vision-only occupancy prediction requires recovering a semantic 3D occupancy field from calibrated surround-view images, where each view provides observations with ambiguous depth along camera rays. Existing methods have progressed from dense structured representations to sparse…

View free PDFSource page
arxivcs.CV2026-07-15

GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

Changqing Zhou, Yueru Luo, Yulan Guo, Bing Wang, Jie Qin, Changhao Chen

Accurate 3D scene understanding is fundamental to embodied intelligence and autonomous driving, where 3D occupancy provides a unified representation of objects, structures, and free space. However, recovering such a complete volumetric representation from visual observations rema…

View free PDFSource page
arxivcs.CV2026-07-24

Learning Adaptive Semantic Gaussian Allocation for 3D Occupancy

Kanglin Ning, Yiran Zhao, Wenrui Li, Houde Quan, Qifan Li, Xingtao Wang, et al.

Semantic 3D Gaussians provide a compact representation for 3D semantic occupancy prediction by rendering semantic primitives into a voxel volume under voxel-wise supervision. Recent methods have improved the modeling ability and efficiency of this representation through more flex…

View free PDFSource page