Highlight
This study reveals that a guideline-anchored retrieval-augmented generation (RAG) large language model configuration significantly surpasses both a baseline GPT-5 and a literature-anchored AI in gynecologic oncology decision support. Key outcomes emphasize the importance of integrating trusted guideline repositories like NCCN to enhance the accuracy and reduce hallucinations in AI-generated clinical recommendations.
Inter-rater reliability was moderate for the primary evaluation metric, supporting the reproducibility of the findings. This benchmarking work precedes clinical integration, underscoring performance differences rather than clinical safety or patient outcome improvements.
Study Background
Gynecologic oncology presents complex diagnostic and therapeutic decision-making challenges due to evolving standards and diverse case presentations. Artificial intelligence (AI), particularly large language models (LLMs), has emerged as a promising tool for augmenting clinical decision support by processing extensive medical information and generating recommendations.
However, LLMs frequently suffer from “hallucinations,” or the generation of inaccurate or unsupported content, limiting their clinical reliability. Integrating authoritative sources such as the National Comprehensive Cancer Network (NCCN) guidelines directly into AI models may mitigate these issues by anchoring outputs in validated knowledge.
Study Design
This study benchmarked three AI configurations in gynecologic oncology decision support using fifty de-identified clinical cases submitted between October and November 2025:
- Baseline GPT-5: A standard large language model without augmented retrieval.
- NCCN-Anchored GPT-5 RAG: A retrieval-augmented generation model that integrates NCCN guidelines directly, enabling retrieval of up-to-date and authoritative knowledge during response generation.
- OpenEvidence: A clinical AI system anchored in literature but without NCCN guideline access at the time of testing.
Three expert gynecologic oncologists scored the AI outputs using a modified Generative Performance Score (mGPS), which ranges from -1 to +1, combining guideline concordance and penalties for hallucinated (erroneous) content. Additional metrics included a Hallucination Penalty component and assessments of readability and rationality. Statistical analyses comprised Wilcoxon signed-rank tests and mixed-effects ordered logistic regression to compare performance across configurations.
Key Findings
The NCCN-anchored GPT-5 RAG configuration delivered the highest mGPS (mean 0.83, SD 0.26), significantly outperforming both the baseline GPT-5 (mean 0.65, SD 0.31) and OpenEvidence (mean 0.70, SD 0.27). Specific statistical comparisons included:
- GPT-RAG vs. baseline GPT-5: W = 189.5, Z = -3.42, P < 0.001, effect size r = 0.49.
- GPT-RAG vs. OpenEvidence: W = 254.0, Z = -2.64, P = 0.008.
- OpenEvidence vs. baseline GPT-5: no statistically significant difference (P = 0.22).
Mixed-effects ordered logistic regression confirmed that GPT-RAG had a substantially higher likelihood of superior mGPS results (odds ratio 3.74; 95% CI, 1.57-8.90).
Inter-rater reliability varied by scoring domain, with an intraclass correlation coefficient (ICC) of 0.49 for the mGPS, indicating moderate agreement among expert scorers. The Hallucination Penalty component showed lower reliability (ICC 0.30), while assessments of readability and rationality had higher consistency (ICC 0.70).
Notably, OpenEvidence subsequently integrated NCCN guidelines on April 27, 2026, externally validating the operational importance of direct guideline anchoring for performance enhancement.
Expert Commentary
This study addresses a critical barrier in applying AI for clinical decision support in gynecologic oncology by directly comparing baseline, literature-anchored, and guideline-anchored models. The clear superiority of NCCN-anchored RAG underscores the necessity of embedding validated, guideline-derived knowledge within LLM architectures to reduce inaccuracies and improve clinical fidelity.
Inter-rater reliability statistics suggest adequate consistency but highlight the inherent complexity in scoring AI-generated clinical recommendations. While the findings are promising, it is crucial to note this benchmark does not equate to established clinical safety or guaranteed improvement in patient outcomes. Implementation studies and prospective evaluations remain necessary.
From a mechanistic perspective, retrieval-augmented generation approaches enable dynamic querying of curated databases during response generation, thereby mitigating knowledge cutoff limitations and contextual errors common in static LLMs. The integration of NCCN guidelines, recognized as a gold standard in oncology, leverages the highest-level structured evidence to support decision-making.
Limitations include the relatively small case sample, restriction to pre-integration scenarios, and the absence of real-world outcome data. The study also illustrates the challenges of balancing AI interpretability, hallucination risk, and timeliness of guideline updates.
Conclusion
This pre-integration benchmark study demonstrates that gynecologic oncology decision support models incorporating direct retrieval and integration of authoritative NCCN guidelines outperform baseline and literature-only AI configurations. The findings highlight the critical operational advantage of guideline anchoring in improving the accuracy and reliability of AI-generated recommendations.
Future research should focus on validating these systems in clinical workflows, assessing impacts on patient outcomes, safety, and clinician acceptance, alongside further enhancing AI transparency and mitigating hallucinations.
Funding and ClinicalTrials.gov
The article does not specify funding sources or registered clinical trials associated with this study.
References
- Dukes D, Yost C, Wang R, et al. Guideline-anchored retrieval-augmented generation outperforms baseline and literature-only configurations in gynecologic oncology decision support: A pre-integration benchmark. Gynecol Oncol. 2026 Jul 16;211:224-230. PMID: 42462288.
- National Comprehensive Cancer Network. NCCN Clinical Practice Guidelines in Oncology. https://www.nccn.org/
- Lee J, Yoon W, Kim S, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems. 2020;33:9459-9474.
- Jha D, Topol EJ. Adapting AI for clinical decision support: opportunities and challenges. N Engl J Med. 2021;384(3):198-202.

