CORTEXA
← Browse
arxivcs.CV2026-07-10

TDSal: Task-Based Top-Down Saliency Prediction Model

Can Mizrakli, Tolga K. Capin

Visual saliency aims to predict the regions of an image most likely to attract human visual attention. While most saliency models assume free-viewing conditions, human attention is often shaped by explicit task goals. In this work, we address task-driven saliency prediction by proposing a model that conditions visual attention on natural-language task descriptions. The model produces task-dependent saliency maps that reflect how attention shifts under different viewing intents. Through quantitative and qualitative analysis, we show that incorporating explicit task semantics enables more faithful modeling of goal-directed visual attention.

View free PDFSource page

Related papers

arxivcs.CV2026-06-30

Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning

Hongyi Lin, Yang Liu, Jinhua Zhao, Xiaobo Qu

Foundation models are increasingly integrated into embodied intelligence systems, but directly assigning them structured prediction tasks requires precise geometric and numerical estimation, where specialized models often remain stronger. This capability mismatch raises a key que…

View free PDFSource page
arxivcs.CVcs.AIcs.RO2026-07-17

DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

Jehun Kang, Jungha Wang, Youngjun Hwang, David Hyunchul Shim

Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature encoders, existing decoding s…

View free PDFSource page
arxivcs.CVcs.AI2026-07-08

ReMoDEx: A Local-to-Global Relevance-Based Model Decision Explainability Framework for large-Scale Image Datasets

Abhay Kumar Pathak, Mrityunjay Chaubey, Manjari Gupta

Deep learning image classifiers achieve strong predictive performance yet remain opaque in how decisions are formed. A model may predict correctly while relying on irrelevant cues, shortcut associations, peripheral structures, or device level artifacts instead of task relevant re…

View free PDFSource page
arxivcs.LGcs.CV2026-07-14

Adversarial Attacks on Online Handwriting using Salience-based Temporal Editing

Yataro Tamura, Brian Kenji Iwana, Jiseok Lee

Deep learning models for online handwriting recognition have been shown effective and are increasingly deployed in practical applications. However, their vulnerability to adversarial attacks is still a challenge. Existing adversarial methods are predominantly designed for image-b…

View free PDFSource page