CORTEXA
← Browse

Han Li

15 papers indexed

arxivcs.CV2026-07-23

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

Han Li, Si Liu, Zehao Huang, Dongxin Lyu, Longfei Xu, Jiahui Fu, et al.

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation of the real world, such as spatial perception and dynamic r…

View free PDFSource page
arxivcs.LGphysics.comp-phphysics.flu-dyn2026-07-21

Thermodynamics-Informed Input Reparameterization for Neural Prediction of Real-Fluid Thermodynamic Properties in Supercritical Combustion

Haoze Zhang, Han Li, Ke Xiao, Yangchen Xu, Runze Mao, Zhi X. Chen

Real-fluid thermodynamic property evaluation is a major computational cost in supercritical combustion simulations. In the enthalpy-based pressure-correction formulation, the closure evaluates temperature T, density $ρ$, and compressibility coefficient $ψ$ from the solver state (…

View free PDFSource page
arxivcs.LG2026-07-12

Policy-Driven CT-Agent: Modeling Phase-Aware Diagnostic Control for Clinically Consistent CT Reasoning

Yanmeng Dong, Han Li, Yujia Li, Jingsong Liu, Xun Ma, Yanzhu Hu, et al.

Computed Tomography (CT) diagnosis often relies on dynamic selection of imaging phases, such as non-contrast, arterial, or venous phases, based on preliminary findings, clinical suspicion, and diagnostic guidelines. This phase-wise decision process is critical for reducing unnece…

View free PDFSource page
arxivcs.SEcs.AI2026-07-10

Failure as a Process: An Anatomy of CLI Coding Agent Trajectories

Xiangxin Zhao, Han Li, Shuaiting Li, Tianyi Zhao, Earl T. Barr, Federica Sarro, et al.

Large language model (LLM) coding agents are increasingly deployed to autonomously perform software engineering tasks in terminal-based environments, making their reliability a growing concern. Existing empirical studies investigate why coding agents fail, yet they largely treat…

View free PDFSource page
arxivcs.CLcs.AIcs.MA2026-07-09

WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search

Xiaoshuai Song, Liancheng Zhang, Kangzhi Zhao, Yutao Zhu, Zhongyuan Wang, Guanting Dong, et al.

Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question answering into complex, deep-and-wide search and research-oriented tasks. A single ReAct-style agent is constrained by one long trajectory and limited context, mak…

View free PDFSource page
arxivcs.CV2026-07-04

Paired Uterine Whole-Slide Images and Pathology Reports for Multimodal Computational Pathology

Han Li, Jingsong Liu, Ayako Ura, Junlin Hou, Zhengyang Xu, Azar Kazemi, et al.

Uterine diseases represent an important category of gynecologic pathology and require accurate histopathological assessment for diagnosis and treatment planning. Whole-slide images (WSI) have enabled the digital transformation of pathology workflows and provided new opportunities…

View free PDFSource page
arxivcs.CV2026-06-30

Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior

Jiahui Fu, Zehao Huang, Han Li, Naiyan Wang, Si Liu

Lane topology reasoning aims to construct a lane graph from onboard sensor observations. Existing methods follow a detection and association paradigm that treats each lane instance independently, leading to geometric inconsistency at connected endpoints and incomplete graphs due…

View free PDFSource page
arxivcs.CV2026-06-30

Towards Voxel Spacing Consistency for Medical Image Segmentation

Xin You, Runze Yang, Minghui Zhang, Hanxiao Zhang, Han Li, Yi Yu, et al.

Volumetric medical image segmentation is essential for both preoperative diagnosis and intraoperative guidance. While recent years have witnessed rapid progress in segmentation architectures, comparatively little attention is paid to the physical voxel spacing of anatomical data.…

View free PDFSource page
arxivcs.CVcs.AI2026-06-26

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

Yankai Yang, Yancheng Long, Bin Wen, Fan Yang, Tingting Gao, Han Li, et al.

Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and differ only in a short time span or a small region, current…

View free PDFSource page
arxivcs.AIcs.CLcs.IR2026-06-25

AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems

Changxin Lao, Fei Pan, Guozhuang Ma, Han Li, Huihuang Lin, Jijun Shi, et al.

Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypothese…

View free PDFSource page
arxivcs.CV2026-06-25

SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

Yankai Yang, Yancheng Long, Wei Chen, Xingyu Lu, Hongyang Wei, Bin Wen, et al.

Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a single whole-image reward, which makes fine-grained editing optimization difficult. We observe that a key obstacle in image editing i…

View free PDFSource page
openalexKnowledge Commons (Lakehead University)

Regulating the Frontier: Model-Level Evidence from the EU AI Act

Han Li

The European Union's Artificial Intelligence Act creates a binding, economy-wide transparency regime directed specifically at providers of general-purpose AI models. Whether such rules alter model-development practice, rather than merely producing legal documentation outside publ…

Also available via: Lakehead University

View free PDFSource page