CORTEXA
← Browse
arxivcs.CV2026-06-25

See & Sniff: Learning Visuo-Olfactory Representations

Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Son Chung, Arda Senocak

While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce SmellNet-V, a scalable visuo-olfactory dataset built on the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows us to synthetically pair smell-only samples with semantically aligned in-the-wild web images, converting a unimodal olfactory dataset into a cross-modal benchmark without costly co-collection. Building on this dataset, we propose See & Sniff, a self-supervised framework that learns joint visuo-olfactory representations via dense local alignment and naturally produces smell saliency maps for spatial grounding of odor sources. We further introduce pixel-level smell localization task and a benchmark for evaluation. Our method surpasses smell-only baselines by 7% in smell classification from smell alone and generalizes to cross-modal retrieval and smell localization, establishing visuo-olfactory learning as a new direction in multimodal perception.

View free PDFSource page

Related papers

arxivcs.CVcs.AIcs.LG2026-07-07

What Images Cannot Say: Language-Guided Olfactory Representation Learning

Eleftherios Tsonis, Xi Wang, Vicky Kalogeiton

Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with electronic-nose measurements, aligning smell signals with images remains challenging because many olfactory cues arise from contextual environment…

View free PDFSource page
arxivcs.CV2026-06-28

CellDETR: A Detection-Guided Framework for Scalable Cell Representation Learning from Histopathology Images

Shikang Zhang, Guojun Li, Yicong Mao, Chulin Sha

Recent advances in pathology foundation models have substantially improved patch and slide level representation learning from whole-slide images (WSIs).However, cell-level representations learning remain underexplored, limiting cell resolved interpretability, biological discovery…

View free PDFSource page
arxivcs.CV2026-07-08

HPR-SAM: Hierarchical Probabilistic Representation Learning for Prompt-free SAM-based Medical Image Segmentation

Yingzhen Hu, Yiheng Zhong, Keying Zhu, Zimu Zhang, Zihan Ye, Sifan Song, et al.

Prompt-free adaptation of the Segment Anything Model (SAM) has emerged as a promising paradigm for automatic medical image segmentation. Existing methods mainly focus on prompt generation, while overlooking that prompt quality is fundamentally constrained by the expressiveness of…

View free PDFSource page
arxivcs.CV2026-07-10

Probing Diffusion Denoising Dynamics for Contrastive Representation Learning

Yasong Dai, Zeeshan Hayder, David Ahmedt-Aristizabal, Hongdong Li

Text-to-image diffusion models exhibit unprecedented generative capability and contain rich intermediate representations that can be useful for discriminative vision tasks. Motivated by this observation, we study a focused question: how can the denoising dynamics of a pretrained…

View free PDFSource page
arxivcs.CV2026-07-14

CGRL: Concept-Guided Pruning and Representation Learning for Whole-Slide Image Classification

Thuc Huynh, Tuan Le, Doanh C. Bui

Weakly supervised whole-slide image (WSI) classification is widely used in computational pathology because slide-level labels are easier to obtain than dense region annotations. Existing multiple instance learning (MIL) methods often aggregate large bags of patch embeddings using…

View free PDFSource page
arxivcs.CVcs.CR2026-07-07

Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning

Zhen Li, Gang Cao, Tian Zhang, Lifang Yu, Shaowei Weng

The rapid advancement of large-scale generative models has accelerated the spread of highly deceptive AI-generated images, making generalized synthetic image detection a critical imperative. Existing forensic networks often struggle with cross-model generalization and realworld d…

View free PDFSource page