This website and third-party tools we use rely on cookies for the best user experience. By selecting "I agree", you agree to cookie usage as described in our Privacy Policy.
1,267 posters, 47 videos, 13 topics, 4 sessions, 853 authors
ePostersLive by SciGen Technologies S.A. All rights reserved.
September 9 - 12, 2026 | George R. Brown Convention Center, Houston, Texas
ALL - 1320
Acute Lymphoblastic Leukemia (ALL)
BACKGROUND
Accurate acute leukemia diagnosis requires integration of morphology, immunophenotyping, cytogenetic/molecular testing, and clinical context. Multimodal large language models (LLMs) can interpret clinical images, but their reliability, calibration, and safety in hematopathology remains uncertain.
OBJECTIVE
To compare three publicly accessible multimodal LLMs for acute leukemia image interpretation using curated peripheral blood (PB) and bone marrow (BM) images.
METHODS
Twenty-two deidentified PB or BM images representing B-ALL NOS (n = 4), B-ALL BCR-ABL-like (n = 3), T-ALL (n = 5), AML-NOS (n = 5), and AML-MRC/AML-MR (n = 5), were selected from the ASH Image Bank. Analyses for myeloid diseases were considered exploratory. Captions and diagnostic labels were withheld. ChatGPT 5.2, Gemini 3.0 Pro, and ChatGPT for Clinicians (Models 1, 2, and 3, respectively) evaluated each image using identical zero-shot, non-iterative prompts. Outputs were scored against the source diagnosis for diagnostic concordance and qualitatively reviewed for morphologic reasoning, differential diagnosis, recommended next steps, anchoring, lineage misassignment, and overconfidence.
RESULTS
Overall diagnostic concordance was 20/22 (91.0%) for Model 1, 12/22 (54.5%) for Model 2, and 19/22 (86.4%) for Model 3. Concordance for Models 1, 2, and 3 was 100%, 75%, and 100% in B-ALL NOS; 100% in all models in B-ALL BCR-ABL-like; and 80%, 20%, and 80% in T-ALL. Exploratory analysis for myeloid diseases showed concordance for Models 1, 2, and 3 of 90%, 50%, and 80%, respectively. All models generally recommended urgent hematology evaluation and confirmatory flow cytometry/cytogenetic/molecular testing. Model 2 more frequently anchored on isolated morphologic features and generated overconfident incorrect diagnoses, whereas Models 1 and 3 more often maintained broader differentials but occasionally misassigned lineage from morphology alone.
CONCLUSIONS
Multimodal LLMs demonstrated variable but potentially useful performance for acute leukemia morphology, with Models 1 and 3 outperforming Model 2 in this curated image set. These findings support further validation of LLMs as educational or preliminary triage tools, including in the settings with limited hematopathology access. Because several tested entities are molecularly or clinically defined, future studies should use larger blinded datasets with standardized prompts, expert adjudication, and integration of immunophenotypic and genomic data.