CORTEXA
← Browse
arxivcs.CV2026-07-09

Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception

Ludovica de Paolis, Marco Baroni, Alessandro Laio, Eugenio Piasini

In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they can exhibit with neural and behavioral data in humans and animals. However, it remains unclear to what extent this alignment persists for visual tasks that extend beyond the canonical object recognition paradigm based on well defined semantic content. In this study, we diverge from the common object-centric view by focusing on another aspect of vision: texture perception. We consider textures of different complexity generated with three different algorithms from the same source images. Using a rank-based statistic, we quantify the information encoded in the internal representations of a CNN and three Vision Transformers (ViTs), and we compare the similarity of these representations to those inferred from human psychophysics data. We find that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representations. Taken together, these results suggest that ViTs may capture more faithfully than CNNs how texture patterns are visually processed by humans, and that the representations of texture stimuli in computational models may be driven by the network architecture.

View free PDFSource page

Related papers

arxivcs.CV2026-06-30

MuSViT: A Foundation Vision Model for Sheet Music Representation

Carlos Penarrubia, Antonio Rios-Vila, Eliseo Fuentes-Martinez, Juan C. Martinez-Sevilla, Francisco J. Castellanos, María Alfaro-Contreras, et al.

Foundation models have transformed vision and language processing by providing rich, reusable representations that transfer across diverse tasks. Sheet music, as a visual encoding of musical language, lacks such a strong domain-specific backbone. We introduce MuSViT (Music Score…

View free PDFSource page
arxivcs.CV2026-07-17

Do Vision Encoders Exhibit Human-like Color Thresholds?

Engy Ehab, Pablo Hernández-Cámara, Nahla Belal, Jesús Malo, Javier Vazquez-Corral, Alexandra Gomez-Villa

Understanding and characterizing human color perception is a longstanding research goal. One of the most traditional approaches is looking for the human color discrimination thresholds, the minimum chromatic differences perceptible to human observers. In recent years, deep neural…

View free PDFSource page
arxivcs.CVcs.AIcs.LG2026-07-03

RADIO1D: Elastic Representations for Condensed Vision Modeling

Greg Heinrich, Mike Ranzinger, Collin McCarthy, Natan Bagrov, Eugene Khvedchenya, Bryan Catanzaro, et al.

This paper challenges the assumption that vision-language models (VLMs) require fixed patch-based 2D vision features. Analyzing fine-tuned vision encoders, we find that representations become increasingly abstract and less spatially coherent during VLM training. Notably, models t…

View free PDFSource page
arxivcs.CV2026-06-26

AI-Generated Image Recognition via Fusion of CNNs and Vision Transformers

Xuan-Bach Mai, Hoang-Minh Nguyen-Huu, Quoc-Nghia Nguyen, Hoang-Tung Vu, Minh-Triet Tran, Trung-Nghia Le

Recent advancements in synthetic data technology have opened a new era where images of remarkable quality are generated, blurring the lines between real-life images and those produced by Artificial Intelligence (AI). This evolution poses a significant challenge to ensuring the re…

View free PDFSource page