CORTEXA
← Browse
arxivcs.CV2026-07-22

ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment

Stefanos Gkikas, Yu Fang, Christian Arzate Cruz, Muhammad Umar Khan, Raul Fernandez Rojas

Automatic pain assessment from facial video remains challenging due to the spatial heterogeneity of pain-related facial cues. This study proposes ReFace, a spatial reorganization pipeline that divides facial input into four spatial quadrants before tokenization, rather than processing the entire face as a single region. Evaluated on the AI4Pain dataset, the proposed approach achieves $56.00\%$ accuracy on the test set using video only, achieving the highest reported accuracy under the fixed AI4Pain benchmark protocol among the compared methods. Notably, the four-quadrant configuration processes the same total pixel budget as the full-face input, yet achieves higher accuracy, suggesting that spatial reorganization can improve performance under the proposed tokenization design. A single quadrant region, processing just one quarter of those pixels, remains competitive at a fraction of the computational cost.

View free PDFSource page

Related papers

arxivcs.CV2026-06-25

TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment

Shuchao Duan, Alan Whone, Hossein Rahmani, Jun Liu, Majid Mirmehdi

Existing facial expression quality assessment (FEQA) methods typically produce only a severity score, without explicitly communicating the observable facial motion evidence that supports the prediction. This limits interpretability and makes it difficult to inspect the basis of m…

View free PDFSource page
arxivcs.CV2026-06-30

Rhythm-Structured Predictive Learning for Remote Photoplethysmography

Ba-Thinh Nguyen, Huu-Dung Nguyen, Thi-Duyen Ngo, Thanh-Ha Le

Remote photoplethysmography (rPPG) estimates physiological signals from facial videos by analyzing subtle pulse induced skin color variations. Despite recent progress, existing self-supervised rPPG methods mainly reconstruct masked pixels or low-level visual representations, whic…

View free PDFSource page
arxivcs.CV2026-07-03

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

Syed Ariff Syed Hesham, Yun Liu, Guolei Sun, Jing Yang, Henghui Ding, Xue Geng, et al.

Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense spatiotemporal tokens whose quadratic self-attention cost makes long-video processing prohibitive. Existing methods address this t…

View free PDFSource page
arxivcs.CV2026-06-28

CellDETR: A Detection-Guided Framework for Scalable Cell Representation Learning from Histopathology Images

Shikang Zhang, Guojun Li, Yicong Mao, Chulin Sha

Recent advances in pathology foundation models have substantially improved patch and slide level representation learning from whole-slide images (WSIs).However, cell-level representations learning remain underexplored, limiting cell resolved interpretability, biological discovery…

View free PDFSource page