Abstract
Automated radiology reporting workflows are often evaluated by model accuracy, but their clinical value depends on how they redistribute work load and residual risk between the model and the radiologist. We quantified the radiologist correction checkpoint in a two-stage speech-to-report workflow, in which draft reports from three automatic speech recognition (ASR) models (GPT-4o-transcribe, Whisper + gpt-oss post-processing, and Whisper) were mapped to a CEUS LI-RADS v2017 category by a rule-based categorizer, across 163 contrast-enhanced ultrasound liver observations. Raw category accuracy in creased with ASR quality (42.3%, 78.5%, 97.5%), but after radiologist correction it converged above 95% across all systems, so accuracy was secured by the checkpoint rather than the model. Correction burden scaled inversely with ASR quality (correction time 11.7 vs 42.9 s; effort 1 vs 4), and clinically-impactful errors capable of changing the category occurred in every system, including the best (18/163). Residual category error after correction was highest where correc tion burden was greatest. Therefore, human checkpoints should be treated as measurable and designable components of automated reporting workflows.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/HAIC26_009.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: https://openreview.net/forum?id=qRBwBWOX4W
BibTex
@InProceedings{HanTae_Human_MICCAISAT2026,
author = { Han, Taewon AND Shin, Jaeseung},
title = { { Human Checkpoints in Automated Speech-to-Report Generation: Radiologist Correction for CEUS LI-RADS Categorization } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17279},
month = {pending},
page = {pending}
}
