Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-and-error process, allowing the LLM to both select modification directions and g…
Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult. Existing scalable RL methods typically assign trajectory-level rewards uniformly across tok…
In cell-free massive multiple-input multiple-output (CF-mMIMO) systems, the canonical uplink local receiver is the local minimum mean square error (LMMSE) receiver with large-scale fading decoding (LSFD) at the central processing unit (CPU). The LSFD coefficients are derived unde…