Highlight
- ChatGPT-4 effectively collected histories of present illness from pediatric emergency department patients/caregivers in a waiting room setting.
- High usability, patient satisfaction, and physician-rated quality, accuracy, and readability were demonstrated.
- AI-generated medical histories showed potential to reduce clinician documentation burden and support triage efficiency.
- No hallucinations were observed in final AI-generated transcripts, indicating safety in preliminary use.
Study Background
The emergency department (ED) environment is often characterized by high patient volumes, critical time constraints, and documentation burdens that can impact clinician workflow and patient experience. Pediatric emergency departments, in particular, face the challenge of efficiently obtaining accurate histories of present illness (HPI) from patients or caregivers, which is foundational to informed clinical decision-making and triage prioritization. Traditional history-taking is time-intensive and may limit patient engagement, with potential consequences for care quality and satisfaction. Advances in artificial intelligence (AI), particularly natural language processing (NLP) models such as ChatGPT-4, provide novel opportunities to automate and streamline early clinical data collection. Using AI to assist in patient history collection directly in the waiting room could augment clinical workflows, improve data quality, and reduce provider burden. However, clinical feasibility, accuracy, and acceptability remain to be established, especially in pediatric emergency settings where patient communication varies widely in complexity and context.
Study Design
This prospective mixed-methods pilot study was conducted in a pediatric emergency department. The study population included 31 patients triaged with Emergency Severity Index (ESI) scores 3 to 5, representing moderate to low acuity cases, appropriate for waiting room encounters. Patients or their caregivers interacted with a Health Insurance Portability and Accountability Act (HIPAA)-compliant instance of ChatGPT-4 to generate a history of present illness. This AI-generated history was collected in parallel but deliberately not used to influence clinical decisions, serving as a proof-of-concept assessment.
Participant experience was quantitatively evaluated using an 11-point Likert scale to rate usability, satisfaction, and perceived quality of the AI-generated summary. Additionally, open-ended responses were subjected to thematic qualitative analysis. Pediatric emergency physicians independently reviewed the AI-generated summaries and rated them on accuracy, completeness, efficiency, readability, and overall satisfaction, also using an 11-point scale. The study thus combined quantitative user and expert assessments with qualitative insights to comprehensively evaluate feasibility and acceptability.
Key Findings
Participants demonstrated very high ratings for usability (median 10, interquartile range [IQR] 8–10) and satisfaction (median 8, IQR 7–10). They also rated the overall quality of the AI-generated history of present illness highly (median 9, IQR 8–10). Qualitative thematic analysis revealed positive commentary focused on the user-friendly interface, perceived objectivity of the AI system, and accuracy of the generated summaries. Participants appeared comfortable engaging directly with the AI tool to communicate clinical information.
Pediatric emergency physicians gave favorable ratings across 5 assessed domains: accuracy, completeness, efficiency, readability, and overall satisfaction. Median scores were 8/10 for accuracy, completeness, efficiency, and overall satisfaction, with the highest median score of 9/10 for readability (IQR 7–9). This suggests that AI-generated histories were not only accurate but also presented in an accessible and concise manner, facilitating rapid clinical assimilation.
Clinicians noted some limitations, including occasional absence of important contextual details such as prior visit features or medical nuances that could influence diagnostic reasoning. Importantly, no hallucinations, defined as AI-generated inaccurate or fabricated information, were observed in final transcripts, underscoring safety in this preliminary context.
Implications for Clinical Practice
The findings support the feasibility of integrating AI-driven medical history collection in pediatric emergency settings, particularly for lower-acuity patients waiting to be seen. Such early retrieval of patient histories may enable clinicians to review organized summaries in advance, inform triage decisions, and potentially decrease clinician documentation workload. This approach can enhance patient engagement by involving caregivers directly through an easy-to-use interface. The favorable ratings for readability and efficiency align with workflow demands important in busy ED environments.
Expert Commentary
This pilot study contributes to the growing body of evidence supporting AI applications in frontline clinical care. The use of ChatGPT-4 represents an advance over earlier rule-based or limited NLP systems, offering flexibility and contextual understanding in patient dialogue. However, as physicians highlighted, AI summaries may miss clinically relevant prior visit features or unique historical nuances, indicating the technology should supplement, not replace, clinician judgment. Attention to completeness and context remains critical.
Future research should explore integration with electronic health records to incorporate longitudinal data, real-time clinical decision support, and scalability to diverse patient populations and higher-acuity cases. Validation studies with larger samples and rigorous clinical outcome measures, along with workflow impact and cost-effectiveness analyses, will help define the clinical role of AI-driven history taking. Data privacy, informed consent, and equitable access also warrant ongoing scrutiny.
Conclusion
The study demonstrates that ChatGPT-4–enabled early retrieval of histories of present illness from pediatric emergency department patients or caregivers is feasible, well accepted, and produces accurate, complete, readable, and efficient summaries. This promising approach has potential to alleviate clinician documentation burden, support streamlined triage processes, and enhance patient engagement in the pediatric emergency context. These pilot findings lay the groundwork for future research to optimize AI implementation strategies and evaluate direct impacts on clinical outcomes and healthcare delivery.
Funding and Registration
No specific funding sources or clinical trial registrations have been reported in the study abstract.
References
Morley-Fletcher A, Raghavan VR, Geanacopoulos AT, Maher M, Waltzman M, Barak Corren Y, Fine AM. A Pilot Study to Evaluate Artificial Intelligence-Driven Early Retrieval of Medical Histories in the Emergency Department. Ann Emerg Med. 2026 Mar 19;88(2):135-143. PMID: 41854578. https://pubmed.ncbi.nlm.nih.gov/41854578/
Additional relevant literature includes:
1. Davenport T, Kalakota R. The potential for artificial intelligence in healthcare. Future Healthc J. 2019 Jun;6(2):94-98.
2. Lin S, Mahoney J, Chan A, et al. Natural Language Processing for Clinical Data Analysis: Systematic Review. JMIR Med Inform. 2022;10(6):e26285.
3. Obermeyer Z, Emanuel EJ. Predicting the Future — Big Data, Machine Learning, and Clinical Medicine. N Engl J Med. 2016 Sep 29;375(13):1216-1219.

