Abstract
Non-imaging medical-AI deployments increasingly take the form of source-grounded conversational systems: an assistant retrieves curated content and generates an answer that must not exceed what its sources support. Their pre-deployment evaluation is dominated by aggregate scores that conflate two guarantees, whether claims are grounded in retrieved sources and whether the response addresses what was asked. The methodology reported here, applied to a non-clinical-grade consumer mental-health assistant outside a regulated pathway, has three components: a fixed 100-query split across 20 risk categories; a grounded-refusal contract fixing a refusal token and a strict-unsupported policy before generation; and a four-axis judge protocol on completeness, grounding, correctness, and safety compliance with confidence intervals. The split exposes a separation an aggregate would mask: grounding is essentially perfect (mean 4.78, 95% CI [4.68, 4.87]; strict-unsupported rate 0.0) and safety compliance high (mean 4.60, [4.48, 4.72]; refusal rate 0.01), yet completeness is substantially lower (mean 3.24, [3.04, 3.43]; 18/100 answers scoring ≤ 2/5). Because the 100 queries are nested in 20 categories of 5, every interval is also recomputed with the category as the resampling unit; the separation is unchanged. The split was frozen before the run but also used during development to harden the grounding contract, so these numbers are a contract check, not a held-out generalisation estimate
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/AMAI_007.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: https://openreview.net/profile?id=%7ERob_Sneiderman1
BibTex
@InProceedings{SneRob_PreDeployment_MICCAISAT2026,
author = { Sneiderman, Robert},
title = { { Pre-Deployment Stress Testing of Source-Grounded Conversational AI in Non-Imaging Mental-Health Contexts: A 100-Query, 20-Risk-Category Methodology with Grounded-Refusal Contracts } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17273},
month = {pending},
page = {pending}
}
