This website and third-party tools we use rely on cookies for the best user experience. By selecting "I agree", you agree to cookie usage as described in our Privacy Policy.
119 posters, 6 topics, 524 authors, 243 institutions
ePostersLive by SciGen Technologies S.A. All rights reserved.
29-30 June, 2026 | QEII Centre, Westminster

245
AI vigilance ??? Post-implementation monitoring, real world performance evaluation, health economic evaluation
Purpose
In recent years, there has been rapid development of artificial intelligence (AI) tools within radiology, including for the detection and characterisation of demyelinating lesions in multiple sclerosis (MS). Although many tools have been clinically evaluated initially, the impact of periodic version updates is unclear without careful re-evaluation. We aimed to evaluate the performance of a MDR Class IIa tool (Pixyl.Neuro.MS) across different versions.
Methods
We performed a retrospective 3-month review of MRI brain studies for the monitoring of patients with MS. The performance of Pixyl.Neuro.MS [versions 2.1.0 and 3.0.0] in processing, detecting, and characterising new lesions on 3D T2-FLAIR MRI sequences was assessed on the same set of studies, with presence of any new lesion taken to indicate active disease.
Results
The performance of Pixyl.Neuro.MS markedly varied between the versions [2.1.0; 3.0.0]. Technical processing performance was reduced [43; 28 cases successfully processed]. Diagnostic performance also varied, with paired comparison revealing similar case-level accuracy [89%; 86%] but comprising markedly decreased sensitivity [100%; 70%] and increased specificity [82%; 94%].
Conclusion
Re-evaluation of Pixyl.Neuro.MS after version update from 2.1.0 to 3.0.0 revealed reduced technical processing performance and altered diagnostic performance. Despite similar overall accuracy, the decreased sensitivity and increased specificity of the newer version may ultimately result in reduced clinical utility, if mainly used in radiological practice as an aid for detecting subtle lesions. These findings may generalise to other AI tools, and we advocate careful post-market surveillance. Utilising a standard set of reference cases for repeated evaluation across versions may be able to reveal performance changes with smaller sample size than would otherwise be necessary.