Link prediction aims to identify potential or future connections within a given graph structure. Position information is essential for link prediction, as it distinguishes homogeneous nodes through their relative relationships, facilitating the accurate capture of structural patt…
PURPOSE: To evaluate whether whole-body PET/CT-derived body composition features are associated with survival in patients with resectable non-small cell lung cancer (NSCLC), using double machine learning to quantify adjusted associations with restricted mean survival time. METHOD…
Objective To develop and validate a dual-plane ultrasound (US)-based deep learning radiomics model for the noninvasive discrimination between parotid pleomorphic adenoma (PA) and Warthin tumor (WT). Methods This retrospective two-center study enrolled 656 patients with pathologic…
Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a result, these systems often suffer from the black-box nature and computational complexity of the underlying models, making it difficu…
Consumer wearables increasingly infer sleep stages from signals including heart rate, accelerometry, and photoplethysmography. However, existing studies often report end-to-end performance under a fixed signal setting, making it difficult to determine whether the observed perform…
Wearable photoplethysmography (PPG) provides continuous heart-rate measurements, but its accuracy degrades under motion. In the ring-platform benchmark, the best supervised baseline reaches 5.33 BPM mean absolute error (MAE) on the overall heart-rate task. In the motion-focused r…
The systemic shift in the energy consumption structure from high-carbon fossil fuels to low-carbon clean energy constitutes a critical pathway toward global climate governance and carbon neutrality. However, this sustainable transition is consistently impeded by deep-seated insti…
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse…
Pretrained vision-language-action (VLA) policies provide strong language-conditioned manipulation knowledge, but they remain largely vision-driven and can struggle once manipulation enters contact states where the scene is occluded, depth is ambiguous, or small force errors push…
Most daily activities are inherently procedural. However, existing evaluations for egocentric video understanding seldom address procedural understanding and largely overlook complex key-step-level reasoning under the widely used video question answering (VQA) paradigm for MLLMs.…
Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficien…
High-frequency communications strongly depend on the line-of-sight (LoS) path, and obstacle blockage can severely degrade the received signal power and achievable rate. Near-field Airy beams with curved trajectories can circumvent obstacles, offering a promising way to alleviate…
The increasing uncertainty from flexible demand and renewable generation has made distributionally robust optimization (DRO) an important tool for robust power system dispatch. DRO relies on forecast scenarios to construct ambiguity sets, but conventional scenario generation pipe…
Globally consistent semantic digital twins require centimeter-accurate and geographically transferable 3D facade segmentation. However, progress in facade parsing is limited by the lack of large-scale, standardized benchmarks for evaluating cross-domain generalization. Existing d…
In real-world deployments, scene text detectors inevitably face distribution shifts beyond the training distribution. Prior work often depends on large-scale scene-text pretraining, yet evaluation under cross-domain changes and real-world imaging degradations remains limited. We…
We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) to improve both the visual quality and instruction-following capability of the Qwen-Image-2.0 diffusion model. To provide reliabl…
The volatility of power grid loads and the uncertainty of distributed energy resources pose challenges to operational economy and safety. Flexible load resource regulation is key to mitigating fluctuations and improving energy efficiency, yet traditional optimization methods have…
High-accuracy spatiotemporal monitoring of surface nitrogen dioxide (NO2) concentrations is essential for air quality management. This study evaluates machine learning-based estimates of near-surface NO2 concentrations using data from the geostationary GEMS instrument and the pol…
Potato holds significant importance as a staple food crop worldwide, particularly in addressing the needs of a growing population. Accurate estimation of the potato Leaf Area Index (LAI) plays a crucial role in predicting crop yield and facilitating precise management practices.…