Highlight
– Baseline performance of GPT-5 in predicting same-day hospital discharge showed modest accuracy, with operational workflow factors as the main source of prediction errors.
– Automated prompt optimization at inference time significantly improved LLM performance, achieving predictive metrics comparable to prior discharge prediction models.
– Expert-led prompt engineering and test-time scaling yielded minimal improvements, emphasizing the value of automated data-driven prompt refinement.
Study Background
Accurate prediction of hospital discharge timing is critical for optimizing bed utilization, coordinating post-acute care, and enhancing patient flow. Traditional predictive models often rely on structured data and manual feature engineering. The emergence of large language models (LLMs), such as GPT-5, offers an opportunity to leverage unstructured clinical documentation with advanced natural language understanding to improve discharge predictions. However, LLMs’ applicability in clinical operational contexts requires rigorous evaluation, as errors may stem not only from medical knowledge gaps but also from understanding hospital workflows.
Study Design
This retrospective cohort study was performed at a tertiary academic medical center between June and December 2025. The study included hospitalized adult inpatients (≥18 years old) admitted during 2024 with lengths of stay ranging from 2 to 14 days. Two independent randomized cohorts were generated: a validation set (n=860) and a test set (n=886). Clinical notes and documentation from the 30 hours preceding a 06:00 index timepoint were inputted into large language models to predict same-day discharge.
LLM performance metrics included F1 score, balanced accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Qualitative error analysis was performed to categorize LLM prediction errors, with particular attention to clinical plausibility and operational workflow understanding. Three inference-time optimization strategies were evaluated: test-time scaling (adjusting prediction thresholds), expert-led prompt engineering (clinician-driven prompt adjustments), and automated prompt optimization (algorithmic selection of optimized prompt configurations).
Key Findings
Baseline GPT-5 with a standard prompt demonstrated an F1 score of 0.48 and a sensitivity of 0.37 on the validation cohort, reflecting modest predictive accuracy. Qualitative analysis revealed that mispredictions predominantly involved understanding and integrating operational workflow factors—for example, discharge delays due to non-clinical issues such as pending administrative procedures or bed availability.
Among the optimization strategies, automated prompt optimization yielded the most substantial performance improvement on the independent test set. This approach enhanced both F1 score and sensitivity compared to the baseline prompt, though it resulted in slightly reduced PPV and specificity, indicating an increased false-positive rate but improved capture of true discharges. In contrast, test-time scaling and expert-led prompt engineering produced minimal gains in performance metrics.
The improved predictive performance of automated prompt optimization brought LLM discharge prediction into alignment with the range of results reported by traditional discharge prediction models, suggesting practical utility in clinical settings.
Expert Commentary
This study underscores the complexity of integrating LLM-based approaches in hospital operational workflows. While LLMs possess extensive medical knowledge, successful application in discharge prediction requires modeling institution-specific and process-driven factors that influence patient flow. Automated prompt optimization can dynamically tailor model inputs to contextual nuances, reducing human bias and improving adaptability.
Study limitations include the retrospective design at a single center, which may restrict generalizability to other hospital settings with differing operational constraints. Further prospective validation and calibration with real-time workflow data would be beneficial. Additionally, balancing improvements in sensitivity with specificity remains a clinical trade-off, particularly where false positives could burden discharge planning resources.
Conclusion
This investigation demonstrates that large language models like GPT-5 can predict same-day hospital discharge with moderate baseline accuracy, mainly limited by challenges in operational workflow comprehension rather than medical knowledge deficiencies. Automated inference-time prompt optimization significantly enhances prediction performance, highlighting its potential to support discharge planning and hospital throughput initiatives. Future work should focus on integrating real-time workflow data and expanding model personalization across diverse clinical environments.
Funding and Clinical Trials
The published study did not report specific funding sources or clinical trial registration details.
References
Spagnoli J, Guzman NM, Desai K, et al. Optimizing Large Language Models for Hospital Discharge Prediction. J Gen Intern Med. 2026 Sep 15. PMID: 42745147. Available from: https://pubmed.ncbi.nlm.nih.gov/42745147/
Nguyen PA, Kabaliuk N, Malin B. Predicting Hospital Discharge: Challenges and Opportunities. J Am Med Inform Assoc. 2023;30(2):233-241.
Rajkomar A, Dean J, Kohane I. Machine Learning in Medicine. N Engl J Med. 2019;380(14):1347-1358.
Sendak MP, D’Arcy J, Kashyap S, et al. A Path for Translation of Machine Learning Products into Healthcare Delivery. EMJ Innov. 2019;3(1):1-5.

