GASTROENTEROLOJİ’YE ÖZGÜ YAPAY ZEKA DİL MODELİNİN KLİNİK PRATİKTE KULLANIMI

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Tıp Fakültesi

Abstract

Bozdoğan Ali B.: Clinical Use of a Gastroenterology-Specific Artificial Intelligence Language Model; Hacettepe University Adult Hospital Residency Thesis, Department of Internal Medicine, Hacettepe University Faculty of Medicine, Ankara, 2026. Objective: This study aimed to evaluate the clinical decision support capacity of GastroGPT, a retrieval-augmented generation (RAG)-based artificial intelligence language model developed specifically for gastroenterology, and to perform a comparative analysis with general-purpose large language models. Materials and Methods: In this retrospective, multicenter comparative model evaluation study, 200 standardized clinical cases selected from the archives of Hacettepe University Faculty of Medicine, Department of Gastroenterology were utilized. Cases were stratified to represent the following subspecialty areas: general gastroenterology (50%), hepatology (26%), pancreatic diseases (9%), inflammatory bowel diseases (6%), gastrointestinal oncology (5%), and endoscopy (4%). A total of eight AI language models were evaluated, comprising four GastroGPT variants (Deeplake, Faiss, Chroma, and 16K) and four general-purpose models (GPT-4, Claude, Bard/Gemini, You.com). Model responses were scored by two independent expert gastroenterologists in a blinded fashion using a 5-point Likert scale across five criteria: diagnostic accuracy, appropriateness of recommended investigations, guideline concordance of treatment recommendations, applicability of patient management strategies, and overall performance. Additionally, hallucination rates, source attribution quality, and response consistency were assessed. Results: GastroGPT-Deeplake achieved the highest overall performance score (53.8±5.9; maximum 60 points), significantly outperforming GPT-4 (50.1±6.8, p=0.028) and Claude (49.3±7.2, p=0.012). GastroGPT variants demonstrated marked superiority over general-purpose models in hallucination rates (4.2–6.8% vs. 8.3– 12.1%, p=0.018) and source attribution quality (62.8–78.3% vs. 23.5–34.7%). In the vector database comparison, Deeplake significantly outperformed Faiss and Chroma (p<0.001). GastroGPT's mean scores across the five evaluation criteria ranged from 4.00 to 4.21, with the highest performance observed in guideline concordance (4.21±0.63). Case complexity had a significant effect on diagnostic accuracy (F=7.361, p=0.0008); diagnostic accuracy scores in high-complexity cases (3.81±0.83) were significantly lower compared to low-complexity cases (4.43±0.70; Cohen's d=0.803). No significant performance differences were observed across disease prevalence categories or subspecialty areas. Inter-rater reliability was excellent (ICC>0.99, Pearson r=0.814, Cronbach's α=0.802). ROC analysis yielded an AUC of 0.908 for the overall performance criterion. Conclusion: GastroGPT-Deeplake demonstrated superiority over general-purpose large language models in both overall performance and safety parameters through domain-specific RAG integration. The advantages of RAG technology in reducing hallucination risk and ensuring source traceability are of critical importance for the safety of medical AI applications. Our findings support the preferential use of domainspecific approaches over general-purpose models in developing AI-assisted clinical decision support systems for gastroenterology practice. Keywords: Artificial intelligence, large language model, gastroenterology, GastroGPT, retrieval-augmented generation, clinical decision support system, hallucination, vector database

Description

Citation

Endorsement

Review

Supplemented By

Referenced By