CORTEXA
← Browse
arxivcs.LG2026-06-28

Reliability, Faithfulness, and the Limits of Post-hoc Explanations of Opaque Scientific Models

Nick Oh, Helen Jin

Post-hoc explanation methods are routinely used to interpret scientific machine learning models, with the deliverable understood to be insight into the phenomenon the model has been trained on. The transition may be taken to be secured once the model is reliable enough and the explanation faithful enough. We argue it is not. Reliability checks that the model's predictions match the phenomenon's outcomes, and faithfulness checks that the explanation matches the model, but neither checks whether the model works as the phenomenon works, which is what a claim about structure requires. The chain can support candidate hypotheses under external corroboration, but it cannot, on its own, support claims about how the phenomenon is in fact structured.

View free PDFSource page

Related papers

arxivcs.CVcs.LG2026-07-02

Rethinking Post-Hoc Calibration in Semantic Segmentation

Tristan Kirscher, Kim-Celine Kahl, Balint Kovacs, Maximilian R. Rokuss, Klaus Maier-Hein, Xavier Coubez, et al.

Reliable confidence estimates are essential in semantic segmentation, especially in safety-critical settings where overconfident errors can mislead downstream decisions. Yet modern segmentation models often remain miscalibrated. Post-hoc calibration offers a practical way to corr…

View free PDFSource page
arxivcs.LGcs.AI2026-07-14

Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty

Sarah Al-Shareeda, Gulcihan Ozdemir, Heung Seok Jeon

Smart-building load forecasters are often trained offline on dense, multivariate, high-frequency data, but deployment may provide only hourly, feature-limited inputs. Missing features must then be reconstructed, and their errors can propagate through the model. If this input unce…

View free PDFSource page
arxivcs.LG2026-07-13

Condition-Stratified Robustness Analysis of Post-Hoc Calibration Methods for Probabilistic Classifiers

Gurdeep Singh Virdee

Post-hoc calibration is widely adopted to correct probability estimates from trained classifiers, yet most evaluations report aggregate performance without testing whether that performance holds across distinct operating conditions within a single dataset. We present a pre-regist…

View free PDFSource page
arxivcs.CRcs.LG2026-06-30

Probe Choice Changes Canary-Memorization Verdicts: Three Post-Hoc Disagreement Case Studies in a Text-Dominant LoRA-Tuned Autoregressive Testbed

Zhichao Fan, Zexin Zhuang, Yanhang Li

We audit a fixed prefix-window mean-NLL memorization probe (K=20) on a Qwen2.5-VL-7B canary testbed and report three post-hoc cases where it disagrees with full-span secret NLL or greedy exact-recall. C3 (false negative, window truncation): damage lands on hex tokens outside K=20…

View free PDFSource page
arxivcs.LGcs.AIcs.CL2026-07-23

Training Large Language Models for Self-Explanation Faithfulness

Yeoktatt Cheah, María Pérez-Ortiz, Noah Y. Siegel, Oana-Maria Camburu

We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects its internal decision-making process. While existing work focuses on evaluating faithfulness or using…

View free PDFSource page