Advanced autonomous collision avoidance for maritime navigation: A reinforcement learning approach with ship dynamics and environmental awareness
Lichao Yang, Jingxian Liu, Qin Zhou, Zhao Liu, Yukuan Wang, Yang Liu, et al.
47 papers indexed
Lichao Yang, Jingxian Liu, Qin Zhou, Zhao Liu, Yukuan Wang, Yang Liu, et al.
Zhe Li, Lin Zhao, Hanbo Ren, Ranmao Yang, Xinxin Li, Zhijiang Zhang, et al.
Gaochao Guo, Yan Sun, Rujun Hong, Yalin Lu, Xingjie Chen, Liming Zhao, et al.
Inhibitor of nuclear factor kappa-B kinase subunit epsilon (IKBKE), a member of the serine/threonine kinase family, is an important oncogene in glioblastoma. IKBKE is involved in the progression of multiple tumours in glioblastoma (GBM), including tumour invasion, migration, and…
Xinna Jiang, Hongquan Jiang, Maojie Zhang, Yang Liu, Xiaolu Feng, Zhan Yq
This study proposes a THz-TDS-based framework for spectral optimisation and polyethylene (PE) ageing identification. First, a filtering-pooling and peak attention network (FPAN) is developed to mitigate water vapour interference and system noise under conventional conditions. By…
Guanghu Xie, Mingxu Li, Shuo Zhang, Yonglong Zhang, Yifan Yang, Yang Liu, et al.
Flow policies can represent multimodal action distributions for robot manipulation, yet a robot must execute one action at each control step. When several proposals are sampled, critic-based ranking makes data collection depend on value estimates over candidate actions that may b…
Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, et al.
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in g…
Yuan Guo, Wen Chen, Yang Liu, Qiong Wu, Weiren Zhu
Integrated sensing and communication (ISAC) is a key technology for future wireless networks, calling for hardware-efficient architectures to jointly support communication and sensing. In this paper, a transmissive reconfigurable intelligent surface (TRIS) transceiver is leverage…
Yuan Guo, Wen Chen, Yang Liu, Kunlun Wang, Zhendong Li, Qiong Wu
In this paper, a novel transmissive reconfigurable intelligent surface (TRIS) transceiver is employed to enable an integrated sensing and communication (ISAC) system supporting both communication and sensing. Under both perfect and imperfect channel state information (CSI), we st…
Shiyuan Piao, Fan Zehui, Yang Liu, Hong Cheng, Juepeng Zheng, Jie Zhou, et al.
Accurate short-term wind power forecasting is essential for grid stability and operational planning, yet remains challenging due to the complex interactions between atmospheric conditions and turbine dynamics. However, existing methods fail to effectively incorporate weather fore…
Yang Liu, Weixing Chen, Xinshuai Song, Tao Pu, Siwen Mo, Yongjie Bai, et al.
Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, semantic verification, and persistent experience across heterogeneous embodiments. We present PhyAgentOS…
Xinhao Cai, Yixuan Sun, Minghang Zheng, Qingchao Chen, Xin Jin, Song-chun Zhu, et al.
Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. While recent neural approaches have achieved impressive visual realism, they typically model motion as a continuous signal and neglect its com…
Ting Lei, Jialin Liu, Zhu Xu, Yuxin Peng, Yang Liu
Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories. While such paradigms achieve strong performance on closed-set benchmarks, they fundamentally entangle interaction understanding wi…
Yang Liu, Yuhao Liu, Yunran Wei
We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents' risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetric…
Mingchao Sun, Luyang Tang, Yu Liu, Xu Yan, Zhan Li, Yunwei Zhang, et al.
We present ABot-3DWorld 0, a universal multimodal 3D world model that turns text, image, and video inputs into high-fidelity, explorable 3D worlds. At the heart of our framework is a unified Spatial Generative Primitive (SGP), a compact tuple of a high-quality panorama and a spat…
Yuliang Liu, Zhang Li, Ziyang Zhang, Shuo Zhang, Qiang Liu, Jiajun Song, et al.
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text…
Jing Liu, Kun Yang, Yan Wang, Dingkang Yang, Xiaoshuai Hao, Wei Zhang, et al.
Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges. Within distributed network environments, Multimodal Large Language Models (MLLMs) serve as…
Wenyuan Wang, Lianyu Hu, Hao Wang, Yang Liu
Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principl…
Yijie Qian, Juncheng Wang, Chao Xu, Huihan Wang, Yuxiang Feng, Yang Liu, et al.
As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality in generated content. However, existing evaluation metrics presume structural correctness, reducing…
Yuxiang Feng, Juncheng Wang, Chao Xu, Wenlong Hou, Huihan Wang, Yijie Qian, et al.
Forecasting the future anatomy of slow-evolving neurodegenerative diseases could enable earlier, more targeted intervention and improve clinical trial design, but it remains challenging because true progression signals are subtle in longitudinal MRI. In this low-signal regime, tr…
Utkarsh A. Mishra, Yongxin Chen, Danfei Xu, Yang Liu, Xi Chen, Jiayuan Mao
Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on robotic action data. We refer to this discrepancy as the video-action generalization gap. In this pape…
Shun Liu, Nan Xi, Yang Liu, Tianyu Luan, Xuan Gong, David Doermann
Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understanding. Extending RIS to endoscopic imagery, however, presents unique challenges, including scarce high-quality annotations and com…
Yang Liu, Zhaokai Luo, Huayi Jin, Ruozhou He, Chenchen Hong, Zhiyong Wang, et al.
Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts increase prefill cost, may exceed context limits,…
Yuqi Chen, Vincent Siu, Yang Liu, Dawn Song, Chenguang Wang
Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily. We investigate whether tool-use decisions have any stable internal representation that can be extracted and manipulated, a que…
Cheng Huang, Jia Zhang, Yi Jiang, Yang Liu, Karanjit Kooner, Yadi Liu, et al.
Glaucoma is a leading cause of irreversible blindness worldwide, yet most automated diagnosis systems rely on opaque deep-learning models that offer little clinical interpretability. We present GlaKG, a biomarker-centric fundus knowledge graph that integrates structural biomarker…
Joongwon Chae, Lihui Luo, Yang Liu, Dongmei Yu, Peiwu Qin, Runming Wang, et al.
Memory-based anomaly detection is attractive because it localizes defects from normal images without training a decoder or synthesizing pseudo anomalies. However, most memory methods still use the memory bank as a nearest-neighbor lookup table: a test patch is treated as normal i…
Chunyu Xiang, Yang Liu, Rui Gu, Wanguo Liu, Jingwei Shi
Background The secondary injury cascade following spinal cord injury (SCI) drives severe inflammation and tissue destruction. Although the natural polyphenol Procyanidin B2 (PCB2) has well-documented neuroprotective properties, its specific therapeutic efficacy in SCI, as well as…
Yunzhong Si, Huiying Xu, Xinzhong Zhu, Yang Liu, Yao Dong, Wenhao Zhang, et al.
Small object detection in Unmanned Aerial Vehicle (UAV) imagery remains challenging under adverse conditions, including complex weather, low illumination, and sensor noise. These challenges mainly stem from severe background clutter, fine-grained detail degradation, and suboptima…
Jiaqi Tang, Shaoyang Zhang, Fandong Zhang, Shu Zhang, Yang Liu, Qingchao Chen
The growing number of medical vision foundation models highlights the need for effective model selection. However, mainstream selection methods rely on exhaustive fine-tuning, which is computationally expensive. Most of the existing Transferability Estimation (TE) metrics are pri…
Ante Wang, Jiaqi Fu, Xuanyi Chen, Ruotian Ma, Zhaopeng Tu, Weizhi Ma, et al.
Thinking has emerged as a critical capability for Large Language Models (LLMs) tackling complex tasks. However, its reactive nature, where reasoning is passively triggered only upon receiving a user response, inevitably introduces latency that compromises conversational fluidity.…
Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu
Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details. We aim to learn representations whose matching is stabl…
Lihui Luo, Joongwon Chae, Ziyan Chen, Yang Liu, Siyi Cheng, Weihan Gao, et al.
Traditional Chinese Medicine (TCM) diagnosis, particularly through tongue inspection, faces persistent challenges in subjectivity and reproducibility. The application of multimodal artificial intelligence to TCM clinical tasks, such as syndrome differentiation and prescription ge…
Zongwu Xie, Yonglong Zhang, Yifan Yang, Yang Liu, Guanghu Xie
Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion. This article presents GAP-GDRNet, a geometry-aware RGB framework built on GDR-Net for a single-target synthetic spacecraft benchmark. The method s…
Yongjie Bai, Hanting Wang, Mingtong Dai, Qijun Zhong, Yang Liu, Liang Lin
General-purpose vision-language-action models benefit from large vision-language priors, but effective manipulation also requires anticipating action-relevant scene changes. Existing world-action models often rely on large generative world models or dense future rollouts, which a…
Yaoqi Guo, Yang Liu, Jie M. Zhang, Yun Ma, Yiling Lou, Zhenpeng Chen
Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches from issue reports and code repositories. Bug reproduction tests (BRTs) are an important building block for such agents and have been shown use…
Ruichen Ma, Xiaoyang Zhang, Jian Bai, Guanchao Qiao, Liwei Meng, Ning Ning, et al.
The performance of deep spiking neural networks (SNNs) often relies on batch normalization (BN). However, the advanced dynamic BN variants used in state-of-the-art models introduce runtime multiplications, which weaken the hardware-efficiency motivation of SNNs. To address this t…
Ying Chen, Jinyue Li, Kun Wang, Qiankun Li, Yang Liu
The Segment Anything Model with Concepts (SAM3) heralds a new paradigm for open-vocabulary segmentation through natural language interaction, offering significant potential for medical image analysis. However, effectively adapting such a powerful vision-language model to the dive…
Hongyi Lin, Yang Liu, Jinhua Zhao, Xiaobo Qu
Foundation models are increasingly integrated into embodied intelligence systems, but directly assigning them structured prediction tasks requires precise geometric and numerical estimation, where specialized models often remain stronger. This capability mismatch raises a key que…
Lianyu Hu, Shengqian Qin, Zeqin Liao, Qing Guo, Liang Wan, Wei Feng, et al.
Chain-of-thought (CoT) reasoning has enabled multi-modal large language models (MLLMs) to tackle complex visual reasoning tasks by generating explicit intermediate reasoning steps in natural language. However, this text-based reasoning paradigm is inherently slow at inference tim…
Shengqi Xu, Guojin Zhong, Yang Liu, Fanjie Wang, Hu Luo, Hanyu Zhou, et al.
Visuo-Tactile policies leveraging optical tactile sensors have shown great promise in contact-rich manipulation. These sensors achieve high spatial resolution and multi-dimensional force sensing by utilizing an internal camera to monitor the deformation of their elastic gel surfa…
World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics. We introduce a controlled diagnostic protocol for studying event-conditioned latent physical structure in p…
Yang Liu, Xuxin Tang, Jiahao Xu, Chris North
Dimension reduction and semantic interaction support image clustering by making similarity structure visible and manipulable. Existing semantic interaction methods encode users' clustering criterion (a user-interpretable semantic dimension, e.g., action, location, or mood) from d…
Xiaomeng Fu, Junfan Lin, Yang Liu, Yaowei Wang, Guanbin Li, Liang Lin, et al.
Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary…
Jie Xu, Zhongnan Ye, Di Wang, Shasha Huang, Yang Liu, Yu Xiang
Urban color plays a fundamental role in shaping the visual character and cultural identity of cities. Yet in many contexts, current practices remain fragmented, with color analysis often disconnected from planning implementation and governance. To address this issue, this study p…
Chengxiang Hu, Yang Liu, Yiyao Ding, Yue Jin, Weiwei Han
Chronic atrophic gastritis (CAG) is a precancerous gastric condition with limited therapeutic interventions, and the mechanisms underlying the benefits of Coptis chinensis Franch. (CCF) remain insufficiently defined. This study employed an integrated computational strategy to cla…
Chunxu Zhang, Yuanshan Zhao, Wude Yang, Liuqian Gao, Wenyu Zhang, Yang Liu, et al.
Accurate cutting of salmon parts and surface defect detection are the key steps to enhance the added value of its processing. At present, mainstream manual inspection methods have low accuracy and efficiency, making it difficult to meet the demands of industrialized production. A…
Yanbin Zhao, Yang Liu, Shuang Gao, Guohua Liu, Zhiqiang Wan, Denghui Hu
This study introduces a novel satellite image digital surface model (DSM) reconstruction framework grounded in deep learning methodology. The proposed framework effectively utilizes a rational polynomial camera (RPC) model to establish the mapping relationship between image coord…
Song Zhang, Xinting Yang, Yizhong Wang, Zhenxi Zhao, Jintao Liu, Yang Liu, et al.
In intensive aquaculture, the number of fish in a shoal can provide valuable input for the development of intelligent production management systems. However, the traditional artificial sampling method is not only time consuming and laborious, but also may put pressure on the fish…