CORTEXA
← Browse

Koyar Afrasyab

2 papers indexed

arxivcs.AI2026-07-21

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

Koyar Afrasyab

Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing -…

View free PDFSource page
arxivcs.AI2026-07-20

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs

Koyar Afrasyab

Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency pr…

View free PDFSource page