CORTEXA
← Browse
arxivcs.CV2026-07-03

Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study

Yingtie Lei, Fangxun Liu, Baicheng Wu, Colin Lee, Ziheng Zhang, Junke Yang, Zhiyuan Tao, Xuyan Huang, Shuheng Wang, William Koran, Kyle Park, Elijah H Buckwalter, Cheng-Hsuan Chiang, Tejas Naik, Daniel Yi, Wei-Lun Chao

Identity recognition (e.g., person, animal re-identification) has traditionally relied heavily on static appearance cues. Yet motion--consistent, individual-specific dynamics--can provide a complementary and potentially more robust signature, especially when appearance is weak or variable. This raises a fundamental question: when identity-specific motion cues are clearly present, to what extent do modern video models use them for recognition? To investigate this question, we conduct a systematic diagnostic study and introduce BALLER120, a controlled benchmark of 120 professional basketball players performing free-throws. By focusing on the same multi-phase action across individuals, BALLER120 reduces action-level variation and identity-correlated acquisition biases, enabling fine-grained analysis of identity-specific kinematic patterns. We find that modern video models can predict identity accurately from RGB videos, but often rely on static appearance cues such as faces and jersey regions, even when informative motion cues are available. Strikingly, when appearance is suppressed through silhouette-only or skeleton-only inputs, the same model architectures shift toward motion micro-patterns (e.g., foot placement and elbow bending). Despite containing less visual information, appearance-suppressed representations achieve competitive accuracy and stronger robustness to appearance shifts. Our qualitative analyses further show that appearance-suppressed models attend to distinctive motion patterns across individuals. Overall, our study demonstrates that identity-specific motion signatures are present, informative, and learnable, but modern video models may overlook them in favor of easier static shortcuts unless appearance cues are explicitly suppressed.

View free PDFSource page

Related papers

arxivcs.CVcs.GR2026-07-01

TrajLoc: Trajectory-Attention Localization for Multi-Object Motion Control

Omer Sela, Inbar Huberman-Spiegelglas, Michael Rotman, Sagie Benaim, Avi Ben-Cohen

Controlling the motion of multiple objects in image-to-video (I2V) generation requires preserving object identities while enforcing adherence to distinct target trajectories. This becomes particularly challenging as the number of objects increases and their paths intersect or occ…

View free PDFSource page
arxivcs.CV2026-07-02

Training-free Controllable Human Motion Generation under Heterogeneous Constraints

Xiaofei Hui, Bo Yan, Haoxuan Qu, Hossein Rahmani, Jun Liu

Training-free controllable motion generation has attracted growing interest for enabling flexible constraint enforcement without constraint-specific training. However, existing training-free methods require constraints to be continuous objective-based with differentiable losses,…

View free PDFSource page
arxivcs.CV2026-07-02

QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers

Kyobin Choo, Youngmin Kim, Hyunkyung Han, Geunrip Park, Chanyoung Kim, Sunyoung Jung, et al.

Video diffusion transformers (DiTs) generate high-fidelity and temporally coherent videos, yet motion control remains implicit, primarily relying on text prompts. As a result, achieving desired motion often requires extensive prompt engineering and repeated resampling. While fine…

View free PDFSource page
arxivcs.CVcs.AI2026-07-11

Identity-Consistent Expression Fields: A Disentangled Neural Radiance Field Framework for Few-Shot Facial Expression Synthesis

Minh Tran

Neural Radiance Fields (NeRF) have enabled photorealistic novel-view synthesis of 3D scenes and, in the facial domain, have been extended to reconstruct and animate 3D face models from a small number of images. However, existing few-shot dynamic NeRF methods for facial expression…

View free PDFSource page
arxivcs.CV2026-07-06

When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions

Yan Zeng, Yusuke Hosoya, Huyen T. T. Tran, Takayuki Okatani

Text-guided diffusion inversion is central to image editing, where an image is mapped to an initial latent and then edited by replaying the denoising process under a modified prompt. In practice, however, inversion is often performed with a lower classifier-free guidance(CFG) sca…

View free PDFSource page
arxivcs.CVcs.GR2026-07-02

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

Long Vu, Tan Ngo, Animesh Karnewar, Amir Habibian, Binh-Son Hua, Hung Bui, et al.

Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approaches still lack precise and unified controllability over both o…

View free PDFSource page