arxivcs.CV2026-07-15
Screening Is Effective for Visual Recognition
Shunya Shimomura, Kazuhiro Hotta
Vision Transformer (ViT) has been widely used as a powerful framework for modeling global dependencies among image patches. However, its core component, self-attention assigns softmax-normalized relative weights to all patches, making it difficult to evaluate the relevance betwee…