This website and third-party tools we use rely on cookies for the best user experience. By selecting "I agree", you agree to cookie usage as described in our Privacy Policy.
119 posters, 6 topics, 524 authors, 243 institutions
ePostersLive by SciGen Technologies S.A. All rights reserved.
29-30 June, 2026 | QEII Centre, Westminster

233
Transforming practice and leadership: pilot or test data on implementation of AI into clinical practice, clinical feedback or patient perspectives
Title
Beyond Dragon Medical Practice Edition: a browser-based generative AI platform for natural-speech radiology dictation — preliminary evaluation
Author
Sanyam Katyal, MD, FRCR — Department of Radiology, University Hospitals of Derby and Burton NHS Foundation Trust, Derby, United Kingdom ()
Keywords
radiology reporting; speech recognition; automatic speech recognition; ASR; Dragon Medical; Dragon Medical Practice Edition; DMPE; large language model; LLM; generative AI; GenAI; dictation; word error rate; WER; reporting time; cognitive load; NASA-TLX; SUS; crossover study; preliminary evaluation; informatics; reporting workflow
Background
Speech recognition is the rate-limiting step of radiology reporting. Published per-report error rates for radiology speech recognition range 4.8–89%; in one prospective audit, 22% of finalised reports contained significant dictation errors while 63% of radiologists believed their personal error rate was below 10%. The cost of recognition error is paid in correction time, particularly for trainees and non-native English speakers. Generative artificial intelligence offers a different architecture: real-time speech-to-text followed by a large language model (LLM) that restores punctuation, repairs medical terminology and assembles report structure automatically.
Hypothesis
A browser-based generative AI dictation platform reduces first-pass word error rate (WER) and total reporting time relative to Dragon Medical Practice Edition (DMPE), with the magnitude of gain inversely related to prior DMPE familiarity.
Methods
Within-subject, two-platform, counterbalanced crossover. Ten radiologists (trainee to consultant; subspecialties spanning abdominal, chest, neuro, MSK, cardiac and general imaging) each dictated 20 standardised cases (CT 8 / MRI 5 / US 2 / radiograph 3 / echocardiography 2; 23–63 words per gold-standard report) on both platforms in counterbalanced order, with a 14-day washout between arms (n = 400 dictations). Participation was remote, completed via a browser-based study application using each participant's own equipment with the same microphone for both arms. Primary endpoint: first-pass WER (Levenshtein, normalised) versus a fixed gold-standard report. Secondary endpoints with paired inferential testing: total time per case, correction time, post-correction WER, per-case satisfaction (1–5) and NASA-TLX frustration. Additional descriptive measures: edit-operation breakdown, System Usability Scale (SUS) and Net Promoter Score (NPS). Analysis used paired Wilcoxon signed-rank across 200 (participant × case) pairs with Bonferroni correction across the five secondary endpoints; effect sizes reported as Cohen's d.
Results
Mean first-pass WER was 8.91 ± 7.05% for DMPE versus 2.20 ± 4.12% for the generative AI platform — a relative reduction of 75.3% (paired Wilcoxon p < 0.001; Cohen's d = 1.04). Mean total reporting time per case was 75.2 seconds for DMPE versus 26.0 seconds for the generative AI platform, a 65.4% reduction (saving of 49 seconds per case). 93% of the time saving was attributable to reduced correction effort rather than faster dictation. The generative AI platform outperformed DMPE in every individual participant (all ten 95% confidence intervals lay entirely below zero) and in all 20 cases. Largest absolute WER reductions were observed in trainees and non-native English speakers, supporting the pre-specified equity hypothesis. NASA-TLX frustration fell from 64 to 29 (out of 100); mean per-case satisfaction rose from 3.02 to 4.54 on a 1–5 scale. SUS mean improved from 3.2 to 4.2; NPS from 2 to 7. Mean correction keystrokes per dictation: 142 (DMPE) versus 21 (generative AI). Spoken-command ratio in DMPE: 17.0% of all spoken tokens were punctuation/formatting commands.
Discussion
The benefit is delivered downstream of recognition by the LLM cleanup pass, which repairs medical-terminology mis-recognitions, restores punctuation and assembles sentence structure automatically. DMPE quality scales steeply with voice-profile training time, advantaging long-tenured power users; the generative AI platform's quality is more uniform across users because cleanup is downstream of recognition, narrowing the gap for trainees and non-native speakers — the cohorts currently disadvantaged by the incumbent. Vocabulary-dense studies (such as MRI lumbar spine reports) remain difficult for both platforms, motivating a planned RadLex-anchored medical-terminology post-processor. The generative AI platform is browser-based, requires no local installation and requires no per-user voice-profile training period — properties intrinsic to the architecture.
Limitations
Small preliminary cohort (n = 10) requiring external multi-centre validation. Levenshtein WER weights a hallucinated finding identically to a misspelled word; adjudicated clinically-significant-error rate is the appropriate primary endpoint for any subsequent safety study. The standardised case set cannot capture the full long tail of subspecialty vocabulary. Single ASR vendor and single LLM tested; generalisability across alternative back-ends is not addressed. Order effects cannot be fully excluded despite counterbalancing and washout. The incumbent arm tested DMPE only; Nuance's cloud-hosted successor product Dragon Medical One (DMO) was not evaluated and the findings reported here apply specifically to DMPE.
Conclusion
A browser-based generative AI dictation platform produced substantially lower first-pass WER (2.20% vs 8.91%) and shorter total reporting time (26.0 vs 75.2 seconds per case) than Dragon Medical Practice Edition under matched conditions, with the dominant mechanism being reduced correction effort. Direction of effect was preserved across all 10 participants and all 20 cases, with the largest gains in trainees and non-native speakers. These preliminary findings motivate a multi-centre prospective crossover pilot with three comparator arms (DMPE, Dragon Medical One, generative AI platform), adjudicated clinically-significant-error rate as the primary safety endpoint and total per-case reporting time as the primary effectiveness endpoint.
References (key)
Disclosures
The author declares no financial relationships with Nuance Communications or any speech-recognition or large-language-model vendor relevant to the work presented. No external funding was received and no institutional sponsorship was provided.