CORTEXA
← Browse
arxivstat.MLcs.LG2026-07-03

A Hierarchy of Policy Learning Problems

Hamsa Bastani, Osbert Bastani, Shihan Chen

Policy learning has received substantial attention with the goal of learning policies from observational data for decision-making. A majority of work in this space has focused on developing algorithms for computing policies that minimize regret compared to the optimal policy. However, in many practical settings, there is insufficient data to obtain low regret. As a result, recent work has shifted attention to alternative objectives, most notably, studying whether it is possible to learn an improving policy that statistically significantly outperforms baseline policies. We argue that there is substantial merit in studying a broader range of policy learning problems. When there is insufficient data to learn an improving policy, there may still be useful questions that can be answered. To this end, we provide a mathematical framework for studying the relationships between policy learning problems. We formalize three problems within our framework: beyond the optimal policy problem and the improving policy problem, we also propose the policy existence problem, which aims to determine if an improving policy exists. Within our framework, we show that the policy existence problem reduces to the improving policy problem, which in turn reduces to the optimal policy problem; these reductions prove that each problem is at least as easy as the next one (in sample complexity). A key question remains: is this hardness strict? We provide partial answers. First, the gap between the optimal policy and improving policy problems is strict. For the improving policy and policy existence problems, we prove that a sublinear polynomial gap exists under natural conditions on improving policy learning algorithms. Thus, we may be able to answer questions about the existence of an improving policy even when we cannot find one. These results highlight the value in studying a broader range of policy learning problems.

View free PDFSource page

Related papers

arxivstat.MLcs.LGmath.OC2026-07-08

Expressivity and Statistical Trade-offs in Diffusion Policy Learning

Viet Vu, Renyuan Xu, Jiacheng Zhang, Yufei Zhang

Diffusion-based policies have recently emerged as powerful policy parameterizations for reinforcement learning, representing state-conditioned action distributions as terminal laws of diffusion processes with parameterized drifts. This terminal-law representation has shown substa…

View free PDFSource page
arxivecon.EMcs.LGmath.STstat.MEstat.ML2026-07-20

Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference

Masahiro Kato, Taka Kato

We propose one-step and two-step methods for policy learning with retrieval-augmented generation (RAG). We formulate RAG-based action selection under the potential outcome framework. In the two-step method, vector search retrieves action-specific neighboring evidence in an embedd…

View free PDFSource page
arxivstat.MLcs.LG2026-07-19

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning

Joseph Lazzaro, Alessio Russo, Aldo Pacchiano

In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the learner's objective is to identify an optimal policy in a Markov Decision Process (MDP) with high con…

View free PDFSource page
arxivcs.LGstat.ML2026-07-02

Role-Aware Neural Convex Divergence Heads for Asymmetric Representation Learning

He Huang, Lu Shen, Yunfeng Huang, Li Qi

Many representation learning problems involve directed relations, such as lexical entailment, sentence entailment, ontology hierarchy, and citation links. Standard Euclidean, cosine, and Mahalanobis heads are symmetric, while generic neural scorers can model directionality but pr…

View free PDFSource page
arxivcs.LGcs.AIq-bio.BMstat.ML2026-07-01

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization

Xuefeng Liu, Mingxuan Cao, Qinan Huang, Thomas Brettin, Rick Stevens, Le Cong

Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such reasoning remains a key open challenge. We study this problem in instruction-based molecular optimization, where answer-only supervi…

View free PDFSource page