Evaluating ChatGPT-4o in Critical Care Board Review: High Accuracy but Risky Multimodal Interpretations

Highlight

ChatGPT-4o demonstrates slightly superior accuracy compared to pooled clinician responses on critical care board-style questions; however, its deficient image interpretation and clinical reasoning raise concerns about safety and reliability in high-stakes contexts. Strong performance in pulmonary and surgical domains contrasts with poor outcomes in critical care ultrasound interpretation. Approximately one-third of AI responses carry a risk of clinical harm due to errors.

Study Background

Critical care medicine demands rapid, evidence-based decision-making incorporating multiple data modalities, including clinical findings, laboratory values, and various imaging studies. Board certification in critical care requires proficiency in interpreting complex multimodal information. Large language models (LLMs) like ChatGPT-4o have shown promise in medical knowledge assessment but have been less evaluated for their ability to interpret multimodal questions that include clinical images—a vital skill in critical care diagnostics and management.

Given the increasing interest in leveraging AI for clinical education and decision support, rigorous evaluation of LLMs’ accuracy, reasoning, and safety in critical care-related clinical scenarios is essential. Incorrect interpretations or inappropriate recommendations from AI tools in this context could exert potential harm if integrated into decision-making or education without critical oversight.

Study Design

This observational study assessed ChatGPT-4o’s performance using 183 multiple-choice questions from the Society of Critical Care Medicine item bank. Questions encompassed diverse critical care domains and included multimodal content—combining text with relevant clinical images such as radiographs and ultrasound scans.

A custom ChatGPT-4o profile was developed through a validated framework designed to represent the conditions of critical care board examinations. Fourteen experienced critical care clinicians, including physicians, advanced practice providers, and pharmacists, reviewed the AI responses for accuracy, image interpretation quality, clinical reasoning, and potential for patient harm.

No interventions were applied beyond AI question answering; the setting simulated an exam environment to benchmark ChatGPT-4o against human clinician performance.

Key Findings

Overall Accuracy: ChatGPT-4o achieved a correct answer rate of 74.9%, statistically exceeding the pooled clinician response rate of 71.1% (p = 0.03). This suggests the model’s strong medical knowledge and comprehension capabilities.

Domain-Specific Performance: Highest accuracy was noted in pulmonary disease (91.7%), surgery and trauma (87.5%), and neurologic disorders (81.8%). In these fields, question comprehension was robust, but accuracy declined when image interpretation was required.

Image Interpretation: AI performance was suboptimal, with correct image interpretation in only 61.7% of cases. Particularly poor results were evident in critical care ultrasound questions, with just 51.1% accuracy. These findings highlight challenges for LLMs in processing visual clinical data.

Clinical Reasoning and Support: ChatGPT-4o demonstrated moderate reasoning quality (68.3%) and adequacy in providing supporting information (66.1%). This indicates that while the model understands question content well, its integration and application of medical knowledge to reasoning may be limited.

Potential for Harm: Importantly, 33.3% of AI-generated answers were linked with potential for clinical harm, primarily due to erroneous image interpretation and inappropriate treatment recommendations. This raises critical safety concerns about relying on LLMs in clinical decision-making or education without human verification.

Expert Commentary

This study is among the first to critically evaluate a state-of-the-art LLM’s ability to handle multimodal critical care questions under exam-like conditions. The superior overall accuracy compared to pooled clinicians demonstrates the impressive capacity of ChatGPT-4o to assimilate vast medical knowledge. However, the AI’s limited proficiency in interpreting clinical images—a cornerstone of critical care practice—reveals an important gap.

Errors in image-based questions and reasoning may relate to inherent limitations in current LLM architectures that primarily focus on text processing and have less training on visual data or the clinical integration thereof. The substantial risk of potentially harmful advice underlines the need for rigorous validation before introducing such models into clinical workflows or educational assessment tools.

Future improvements may involve multimodal training paradigms explicitly designed to integrate clinical imaging and diagnostics with text-based reasoning. Until then, clinicians and educators should approach LLM outputs cautiously, ensuring human oversight to mitigate risk.

Conclusion

ChatGPT-4o demonstrates promising accuracy surpassing average clinician performance on multimodal critical care board questions but falls short in critical domains of image interpretation and clinical reasoning. These deficiencies contribute to a significant proportion of potentially unsafe recommendations, underscoring a need for cautious application in clinical settings.

While LLMs hold potential as educational adjuncts or decision support tools, their current limitations necessitate expert human review. Further development integrating visual data processing and complex clinical reasoning will be essential to harness these technologies safely and effectively in critical care medicine.

Funding and Clinical Trials

The cited study was published without disclosed funding sources and did not report any interventional clinical trials.

References

1. Sethi I, Khan S, Lyons PG, et al. Large Language Models Provide Accurate but Potentially Unsafe Answers to Multimodal Critical Care Medicine Board Review Questions. Crit Care Med. 2026;54(9):2245-2255. PMID: 42307255.
2. Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44-56.
3. Rajpurkar P, Chen E, Banerjee O, Topol EJ. AI in health and medicine. Nat Med. 2022;28(1):31-38.
4. Langlotz CP, Allen B, Erickson BJ, et al. A roadmap for foundational research on artificial intelligence in medical imaging: From the 2018 NIH/RSNA/ACR/The Academy Workshop. Radiology. 2019;291(3):781-791.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply