Person re-identification (ReID) serves as a critical component in intelligent surveillance systems, aiming to match identities across disjoint camera networks. While traditional methods primarily rely on single-modal RGB imagery, they are often constrained by environmental challe…
Accurate 3D scene understanding is fundamental to embodied intelligence and autonomous driving, where 3D occupancy provides a unified representation of objects, structures, and free space. However, recovering such a complete volumetric representation from visual observations rema…
The geometric size and regularity of detonation cells are key physical parameters for characterizing detonation waves. Traditional manual measurement of soot foils is time-consuming and subjective, while existing computer vision techniques often exhibit poor generalization on rea…
Multi-label image classification (MLIC), a fundamental task in computer vision, focuses on identifying multiple objects or concepts within an image, underpinning numerous read-world applications, such as autonomous driving, disease diagnosis, recommendation system, and mobile ser…
Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from trajectory synthesis. This approach suffers from candidate dependence, heavy computational overhead, and inconsistencies between sampled actions…
World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, existing approaches focus largely on model architecture design, and how a WAM can efficiently learn better world representations…
In response to the limitations of existing evaluation methods for gas well types in tight sandstone gas reservoirs, characterized by low indicator dimensions and a reliance on traditional methods with low prediction accuracy, therefore, a novel approach based on a two-dimensional…