CORTEXA
← Browse
arxivcs.LGstat.AP2026-06-26

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha

AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains. When uncalibrated AI evaluations are used for model ranking, item scoring, or population-level quality reporting, these biases can directly distort downstream decisions. We propose BACON, a four-stage pipeline that combines budgeted human calibration with multiple AI-judge outputs to produce more accurate annotations. BACON constructs full-coverage auxiliary features for every item, including multi-judge scores, token-level uncertainty statistics, and contextual embeddings. It then collects human labels for a small sampled subset and trains a cross-fitted outcome model to generate calibrated item-level surrogate predictions. These predictions support two use cases: population-level estimation of summary metrics, such as means or quantiles, using an augmented estimating-equation estimator with valid confidence intervals; and individual-level surrogate scoring for item ranking and annotation. BACON treats AI judges as auxiliary measurements rather than ground truth: human labels provide the calibration anchor, while AI-derived signals improve efficiency. Across diverse tasks, domains, and labeling budgets, BACON improves predictive accuracy and ranking consistency, and reduces bias and variance relative to raw AI outputs and purely human-label-based methods. These results show that BACON offers a practical, statistically grounded framework for scalable evaluation with limited human annotation.

View free PDFSource page

Related papers

arxivstat.APcs.LG2026-07-01

Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule

Ahmed Cherif

Pretrained generative foundation models cast forecasting as conditional generation from a learned predictive distribution and forecast unseen series zero-shot. We establish three results that turn their reported success into an actionable, mechanistic understanding. First (a posi…

View free PDFSource page
arxivstat.APcs.AIcs.LG2026-07-18

Optimizing Clinical Trial Protocols Using EHR-Derived Heterogeneous Treatment Effects

Xiaodi Li, Munhuwan Lee, Pengyang Li, Xiaoke Liu, Jose K. James, Patricia A. Pellikka, et al.

Traditional randomized trials often obscure clinically meaningful heterogeneity in treatment response by focusing on average effects. Leveraging real-world data to emulate clinical trials and estimate heterogeneous treatment effects (HTEs) offers a promising path toward more prec…

View free PDFSource page
arxiveess.IVcs.CVcs.LGphysics.med-phstat.AP2026-07-09

ConRad: Efficient Conformal Prediction for Radiomics

Matt Y. Cheung, Ashok Veeraraghavan, Guha Balakrishnan

Radiomic features derived from medical images and segmentation masks are used to support decision making in clinical imaging pipelines. In practice, these features are often computed from predicted masks, but segmentation models can be overconfident or poorly calibrated, making d…

View free PDFSource page
arxivstat.APcs.LGmath.OCstat.MLstat.OT2026-07-16

Proactive Inpatient Bed Requests for Emergency Department Admissions

QIan Cheng, Nilay Tanik Argon, Aniruddhan Ganesaraman, Serhan Ziya

Emergency department (ED) boarding occurs when admitted patients remain in the ED while awaiting inpatient beds. Boarding is a major driver of ED crowding and has been associated with poor patient outcomes. We propose a framework to help EDs reduce boarding time and length of sta…

View free PDFSource page