arxivcs.AI2026-07-21
Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety
Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing -…