CORTEXA
← Browse
arxivcs.CV2026-06-28

Zero-Gated Language-conditioned Human Motion Prediction

Guanhui Qiao, Lu Zhou, Ding Jiang, Jinqiao Wang

Pose histories provide the core kinematic evidence for 3D human motion prediction, but they lack explicit high-level semantic guidance. This paper introduces ZGL, a lightweight language-conditioned predictor that uses captions of the observed motion as a semantic prior while preserving a strong motion backbone as the main source of dynamics. We render only the observed poses, generate a one-sentence description with a vision-language model, encode the caption with a frozen CLIP-L text tower, and project it into a small set of conditioning tokens. These tokens are injected into a DCT-based spatial-temporal Transformer by compact crossattention adapters with zero gates: each adapter output is multiplied by a learnable gate initialized to zero, so the full network is numerically identical to the pose-only baseline at initialization and can learn to use language only when it reduces prediction error. On Human3.6M, ZGL improves overall MPJPE over representative motion-prediction baselines in our comparison. Results on CMUMocap further show that compact caption conditioning transfers to a second benchmark and provides a practical semantic cue for 3D human motion prediction.

View free PDFSource page

Related papers

arxivcs.CV2026-06-26

Long-Term Prediction of Local and Global Human Motion with Occlusion Recovery

Qiaoyue Yang, Sven Heutger, Christopher Niemann, Magnus Jung, Ayoub Al-Hamadi, Sven Wachsmuth

Human motion describes the three-dimensional full-body movement of a person. Anticipating such motion holds significant relevance across a wide range of application domains such as human-robot interaction, autonomous driving, animation, and healthcare. In recent research, spatial…

View free PDFSource page
arxivcs.CV2026-06-25

Beyond MoCap: Scaling Motion Tokenizers with Synthetic Human Motion for Generative Modeling

Yiwen Yan, Wanning He, Yu-Wing Tai

Human motion generation models are fundamentally constrained by the limited diversity of motion capture datasets, which predominantly contain common, repetitive actions and fail to cover the long tail of complex human movements, resulting in a restricted motion vocabulary in lear…

View free PDFSource page
arxivcs.ROcs.CVcs.LG2026-07-02

LIME: Learning Intent-aware Camera Motion from Egocentric Video

Boyang Sun, Jiajie Li, Yung-Hsu Yang, Chenyangguang Zhang, Tim Engelbracht, Sunghwan Hong, et al.

Autonomous robots often need to move their camera before they can act: to inspect an object, reveal an occluded region, or obtain a view that responds to a user's intent. While vision-language navigation translates instructions to base motion and vision-language-action policies m…

View free PDFSource page
arxivcs.CV2026-06-29

Scene-aware Prediction of Diverse Human Movement Goals

Qiaoyue Yang, Amadeus Weber, Magnus Jung, Ayoub AI-Hamadi, Sven Wachsmuth

Anticipation of human behaviours facilitates autonomous systems in proactive planning. Human behaviour could be stochastic due to varying goals. Human goals typically guide their own movement and could therefore help to predict the human trajectory and human motion in the long-te…

View free PDFSource page