CORTEXA
← Browse
arxivcs.CV2026-07-06

Repurposing CLIP to Localize at Pixel Level

Jiaxiang Fang, Shiqiang Ma, Jing Wang, Siyu Chen, Fei Guo, Shengfeng He

Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this capability to pixel-level dense prediction poses challenges due to global feature biases. In this paper, we introduce CLIPix, a simple yet effective framework that repurposes CLIP to perform pixel-level localization. By tracing back CLIP's classification process, CLIPix identifies object-specific attentive regions and repurposes them as pixel-level localization cues. To address noise introduced by global biases, we propose a Noise-Resistant Correction strategy, refining these cues for more precise segmentation. Additionally, we introduce a Localization Embedding strategy to integrate both localization and enriched detail information, enabling accurate, high-resolution segmentation. Our approach preserves CLIP's generalization strength and unlocks its potential for segmenting arbitrary objects. Extensive experiments on the PASCAL and COCO datasets demonstrate that CLIPix achieves state-of-the-art performance, underscoring its effectiveness.

View free PDFSource page

Related papers

arxivcs.CVcs.AI2026-07-20

Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs

Yi Tang, Xinyi Shang, Jiacheng Cui, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, et al.

Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalizatio…

View free PDFSource page
arxivcs.CVcs.AI2026-07-21

OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation

Shuimu Chen, Jing Jin, Nan Su, Hongbo Xu, Zebang Cheng, Wenming Yang, et al.

Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable defect reasoning. However, current LVLM-based IAD methods still struggle to produce precise pixel-level an…

View free PDFSource page
arxivcs.CV2026-07-16

SwinAD: Multi-stage feature reconstruction for unsupervised industrial anomaly detection

Huong Ninh, Chien Thai, Mai Xuan Trang, Vu-Minh Le, Thanh Ha Le, Long Tran

Industrial anomaly detection aims to identify and localize defective regions without relying on exhaustive annotations of all possible defect types. Although recent unsupervised methods have achieved strong performance, most are primarily designed for single-class settings and of…

View free PDFSource page
arxivcs.CV2026-07-01

GenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language Models

Hongkuan Zhou, Tristan Rehm, Nadeem Nazer, Lavdim Halilaj, Jingcheng Wu, Steffen Staab

Industrial inspection requires more than binary anomaly detection: a practical system should determine whether an anomaly exists, localize the defective region, identify the defect type, and provide interpretable visual evidence. Existing CLIP-based methods detect and localize an…

View free PDFSource page
arxivcs.CV2026-07-06

Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images

Weikang Wang, Tobias Weißberg, Florian Bernard

While various works address reflective symmetry understanding in 3D data and images, pixel-level semantic left-right prediction of in-the-wild images remains challenging, due to certain difficulties including the lack of 3D information, occlusion, object pose variation, partialit…

View free PDFSource page
arxivcs.CV2026-07-05

Event Detection in Videos: A Framework for the Development of New Methods

Anastasia Zakharova, Thierry Bouwmans, Anthony Cioppa, Adrien Deliège, Antonio Greco, Anaïs Halin, et al.

Event detection tasks in videos, the most important aspect of video surveillance, aim to detect events either at the pixel-level, frame-level, or clip-level. Plenty of methods intended for event detection in different environments, for various applications, and within different a…

View free PDFSource page