This website and third-party tools we use rely on cookies for the best user experience. By selecting "I agree", you agree to cookie usage as described in our Privacy Policy.
119 posters, 6 topics, 524 authors, 243 institutions
ePostersLive by SciGen Technologies S.A. All rights reserved.
29-30 June, 2026 | QEII Centre, Westminster

154
AI vigilance – Post-implementation monitoring, real world performance evaluation, health economic evaluation
From pilot to roll-out: a real-world model for AI vigilance and continuous monitoring in fracture detection
Fox Aoife, Hughes Helen, Strappelli Gian-Marco, Clapp Nicola, Thorndyke Steve, Lee Graham, Lewis-Towler Kelly, Harding Elaine, David Sarojini | Lewisham and Greenwich NHS Trust
Background: Missed fractures affect 3–10% of ED cases, causing patient harm and delayed care. Radiology reporting delays created a gap between image acquisition and formal diagnosis. AI-assisted fracture detection was deployed as a point-of-care decision-support tool. Evaluation and deployment were structured using BS 30440:2023, applied prospectively — not as a retrospective compliance exercise.
Phased implementation: Four phases were undertaken — retrospective validation, adult ED deployment, paediatric extension, and organisation-wide rollout — underpinned by continuous governance including an AI Governance Board, hazard log, Clinical Safety Case, weekly MDT review, discrepancy logging, and RAID log.
AI output: AI fracture tool output on a paediatric wrist fracture
Deployment volume: 7,620 pilot studies (April–August 2025) and 30,394 roll-out studies (August 2025–March 2026), comprising 25,590 adult and 4,804 paediatric studies. Fracture detection rate: 22.5% adults, 24.9% paediatrics.
Diagnostic performance — adult (n=1,058): Sensitivity 99%, Specificity 94%, NPV 99.3%, PPV 89.7%. Exceeded manufacturer benchmarks. Age-related specificity decline in elderly female patients identified through post-deployment audit, not pre-deployment testing.
Diagnostic performance — paediatrics (n=237): Sensitivity 98.1%, Specificity 97.3%, NPV 97.4%, PPV 91.4%. Dedicated paediatric audit conducted prior to extension.
Hip deep dive (n=203, two sites): Sensitivity 90%, Specificity 93.1%, NPV 98.2%, PPV 91.4%. Site comparison following model version update: UHL (pre-update) sensitivity 94.7%, specificity 88.1%; QEH (post-update) sensitivity 81.8%, specificity 97.8%. Ten percent of patients underwent secondary imaging; 85% had discordant AI outputs. Both periprosthetic fractures were missed — mandatory senior clinician and radiologist review now implemented.
Challenges and mitigations: Hip specificity lower than benchmarks (low-confidence findings) — training and threshold recalibration. Age-related specificity decline in elderly female patients — prospective subgroup monitoring. Multi-view inconsistency risk — mandatory multi-view review and software update. Periprosthetic fractures missed (n=2) — senior review mandated, future variant training planned. AI alert not acted upon (3 cases) — human factors training reinforced. Discrepancy reporting gap — targeted refresher training. Turnaround concerns at 3 months — escalated to MDT and resolved with vendor. Dislocation detection misclassification risk — disabled pre-deployment.
Missed fractures: 20 pre-deployment (July–September 2024) versus 11 post-deployment (July–September 2025), a 45% reduction. Descriptive comparison only. Three post-deployment misses had AI alert not acted upon.
Radiology reporting turnaround time: 13 hours 15 minutes (2024 mean) reduced to 11 hours 13 minutes (2025 mean), an improvement of approximately 2 hours. Observational finding — causal attribution cannot be confirmed.
Health economics: Estimated annual saving of £50,758, comprising missed fractures (£3,942), return to ED (£40,976), and radiologist time (£5,840). Preliminary analysis; litigation costs excluded.
User sentiment — adoption trajectory (n=97, 5 waves): Pre-rollout (n=33): cautious optimism, 74% comfortable trusting AI. Post-rollout (n=7): early positive reception, generally easy to use. Three months (n=28): integrating into workflow, 68% positive workflow impact. Seven months (n=24): embedded into routine use, 79% easy to use, 79% regular use. New sites (n=5): early adoption phase, broadly positive with variable confidence. Trust concerns around overcalling sustained throughout all waves — reflecting healthy critical appraisal, not dissatisfaction.
Patient sentiment (n=25, preliminary): 47% comfortable with AI in healthcare. 50% agree AI as additional tool makes care safer. 74% trust clinicians more than AI. Most common feeling: curiosity.
Key conclusions: High diagnostic performance achieved across adult and paediatric populations in real-world NHS deployment. Safety signals emerged through structured monitoring — not pre-deployment testing. Governance framework enabled iterative improvement at model and workflow level. User adoption: cautious optimism to embedded routine use over 7 months. Successfully scaled from single-department pilot to trust-wide deployment. Estimated £50,758 annual cost saving.
Good outcomes followed from good process, not good technology alone.