This website and third-party tools we use rely on cookies for the best user experience. By selecting "I agree", you agree to cookie usage as described in our Privacy Policy.
1,267 posters, 47 videos, 13 topics, 4 sessions, 853 authors
ePostersLive by SciGen Technologies S.A. All rights reserved.
September 9 - 12, 2026 | George R. Brown Convention Center, Houston, Texas
AML - 1636
Acute Myeloid Leukemia (AML)
CLINICAL FIDELITY OF LARGE LANGUAGE MODELS IN ACUTE MYELOID LEUKEMIA: THE IMPACT OF GUIDELINE PROMPTS
Meenal Gehlawat, MD, Maria Alejandra Molina-Rodriguez, MD, Ayush Sharma, Saksham Gakhar PhD, Dani Castillo, MD
BACKGROUND:
Large language models (LLMs) are artificial intelligence (AI) systems capable of generating human-like medical responses that are increasingly used to support clinical decision-making. Accuracy of LLMs in diagnosis and management of acute myeloid leukemia (AML), a rapidly evolving molecularly complex disease, has not been well studied. This study evaluates the efficacy of five leading LLMs and examines the impact of incorporating up-to-date clinical guidelines into LLM prompts to enhance clinical decision-making.
METHODS:
We developed 50 standardized clinical vignettes covering key AML domains, including diagnosis, cytogenetic and risk classification, pretreatment evaluation, targeted therapy selection, adverse effects, and supportive care. We tested five LLMs: ChatGPT GPT 5.5., Claude Sonet 4.5, DeepSeek V3, Gemini 3 Pro, and Meta AI Muse Spark. Each model received a baseline prompt and a guideline-augmented prompt requiring reference to National Comprehensive Cancer Network (NCCN) and European Society for Medical Oncology (ESMO) guidelines. Blinded expert reviewers scored all responses using a standardized answer key (0=incorrect, 1=partially correct, 2=correct). The maximum possible score was 100. Responses were scored by a single expert reviewer per model; absence of independent dual-scoring is a study limitation.
RESULTS:
Guideline augmentation improved accuracy across all five models. Mean accuracy increased from ~73% at baseline to ~88% with guideline prompts.
Individual models showed different improvement patterns. Highest baseline accuracy was identified in responses from Claude at 80% baseline and 5 points gain with guideline prompting. Accuracy for DeepSeek was lowest with approximately 65% at baseline but improved the most with guidelines. ChatGPT, Gemini, and Meta AI showed gains of approximately 10%, 13%, and 17% respectively.
Domain-specific analysis revealed the largest improvements in minimal residual disease assessment and targeted therapy selection (up to 18% gain), while APL management showed smaller improvements. Knowledge drift related to classification systems were identified across all models.
CONCLUSIONS:
LLMs demonstrate impressive medical knowledge, but they remain susceptible to hallucinations and knowledge drift. This necessitates a shift from knowledge-retrieval to guideline-adherence. Our study highlights a framework where mandatory guideline integration serves as the cornerstone for safe LLM application within oncology. Future studies should validate these findings in clinical settings and evaluate mechanisms for automated guideline updating.