This website and third-party tools we use rely on cookies for the best user experience. By selecting "I agree", you agree to cookie usage as described in our Privacy Policy.
119 posters, 6 topics, 524 authors, 243 institutions
ePostersLive by SciGen Technologies S.A. All rights reserved.
29-30 June, 2026 | QEII Centre, Westminster

202
Maximilian Russe, Anna Fink, Carl Simon, Kai K?�stingsch?�fer, Fabian Bamberg, Alexander Rau
University Hospital Freiburg , Department of Diagnostic and Interventional Radiology, University Medical Center Freiburg, Freiburg, Germany, Department of Neuroradiology, Medical Center, University of Freiburg; Faculty of Medicine, University of Freiburg, Freiburg, German
AI vigilance ??? Post-implementation monitoring, real world performance evaluation, health economic evaluation
Title
AI-RADS: A Framework for Case-Level Assessment of Radiology AI Output
Subtitle
Development and multireader evaluation of a structured system for AI reliability, clinical utility, and report communication
Authors
Maximilian F. Russe¹, Anna Fink¹, Carl P. Simon¹, Kai Kästingschäfer¹, Fabian Bamberg¹, Alexander Rau²
Affiliations
¹Department of Diagnostic and Interventional Radiology
²Department of Neuroradiology
Medical Center - University of Freiburg, Faculty of Medicine, University of Freiburg, Germany
Take-home message
From AI output to accountable action: AI-RADS 1-5: integrate, edit, verify, override, or suppress. Structured for evolving regulatory requirements.
Out now as Publication
Russe MF, Fink A, Simon CP, et al. AI-RADS: A Framework for Assessment of Artificial Intelligence Output in Radiology - Development and Multireader Evaluation. Investigative Radiology. Online ahead of print 2026. doi:10.1097/RLI.0000000000001272
Why AI-RADS?
Radiology AI is entering routine workflows across detection, segmentation, quantification, classification, and generative tasks. However, most systems still lack a standardized case-level method to document whether an individual output was reliable, corrected, overridden, or suppressed. AI errors, image-quality limitations, workflow failures, and radiologist overrides may remain poorly traceable. AI-RADS addresses this need by linking each AI output to a structured reliability category and a corresponding clinical action.
Regulatory need
Real-world deployment of radiology AI requires traceable oversight beyond initial validation. AI-RADS creates auditable case-level documentation. This supports privacy- and regulation-conscious governance, including GDPR/HIPAA-aware documentation workflows, EU AI Act-aligned post-market monitoring, EU MDR surveillance concepts for medical-device software, and US trustworthy-AI / FDA-style real-world performance monitoring.
Objective
To develop and evaluate AI-RADS as a structured case-level framework for assessing radiology AI output reliability, clinical utility, and required reporting action. The study tested whether radiologists can reproducibly apply AI-RADS across representative image-based and generative AI tasks.
The AI-RADS framework
AI-RADS 1: Excellent / optimal utility
Integrate into workflow or report.
AI-RADS 2: Highly accurate / high utility
Use with minor confirmation or edits.
AI-RADS 3: Acceptable / modification needed
Verify and modify before use.
AI-RADS 4: Unreliable / limited utility
Override critical sections.
AI-RADS 5: Unreliable / no clinical utility
Suppress from report or display.
Task labels
[DET] Detection
[SEG] Segmentation
[CLS] Classification
[QNT] Quantification
[GEN] Generative AI
Context modifiers
[IL] Image limitation
[SC] Superior confirmation
[WL] Workflow limitation
Example: AI-RADS 4 [SEG, IL]
Study design
7 AI applications
350 total cases
5 radiologist readers
Each case was assessed using:
AI-RADS category + modifiers + correctness rating
Analysis included:
Agreement and usability analysis
Case distribution:
200 image-based AI outputs
150 generative AI outputs
Image-based tasks:
Fracture detection, chest radiograph pathology detection, lung nodule detection, and stroke MRI segmentation/classification.
Generative tasks:
CT protocol and contrast-phase selection, Spinal Instability Neoplastic Score assignment, and patient-friendly report generation.
Assessment endpoints
Interreader agreement
Alignment with correctness
Modifier use
Usability and workload
Key results
AI-RADS showed high reproducibility across image-based and generative radiology AI tasks and standardizes documentation of radiologist oversight. Structured case-level ratings support safer workflow integration, quality assurance, audit trails, post-market monitoring, and regulation-conscious AI governance.
Overall Krippendorff’s alpha: 0.89
Image-based tasks Krippendorff’s alpha: 0.87
Generative tasks Krippendorff’s alpha: 0.93
Reader-rated correct AI outputs mapped to AI-RADS 1-2: 94.8%
AI-RADS ratings unanimous or within one category: 88.3%
Median System Usability Scale / NASA-TLX: 87.5 / 2
Conclusion and clinical interpretation
Conclusion text is visible as a heading in the extracted poster text, but the body text was not included in the parsed extraction.