GASTROENTEROLOJİ’YE ÖZGÜ YAPAY ZEKA DİL MODELİNİN KLİNİK PRATİKTE KULLANIMI
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Tıp Fakültesi
Abstract
Bozdoğan Ali B.: Clinical Use of a Gastroenterology-Specific Artificial
Intelligence Language Model; Hacettepe University Adult Hospital Residency
Thesis, Department of Internal Medicine, Hacettepe University Faculty of
Medicine, Ankara, 2026.
Objective: This study aimed to evaluate the clinical decision support capacity of
GastroGPT, a retrieval-augmented generation (RAG)-based artificial intelligence
language model developed specifically for gastroenterology, and to perform a
comparative analysis with general-purpose large language models.
Materials and Methods: In this retrospective, multicenter comparative model
evaluation study, 200 standardized clinical cases selected from the archives of
Hacettepe University Faculty of Medicine, Department of Gastroenterology were
utilized. Cases were stratified to represent the following subspecialty areas: general
gastroenterology (50%), hepatology (26%), pancreatic diseases (9%), inflammatory
bowel diseases (6%), gastrointestinal oncology (5%), and endoscopy (4%). A total of
eight AI language models were evaluated, comprising four GastroGPT variants
(Deeplake, Faiss, Chroma, and 16K) and four general-purpose models (GPT-4,
Claude, Bard/Gemini, You.com). Model responses were scored by two independent
expert gastroenterologists in a blinded fashion using a 5-point Likert scale across five
criteria: diagnostic accuracy, appropriateness of recommended investigations,
guideline concordance of treatment recommendations, applicability of patient
management strategies, and overall performance. Additionally, hallucination rates,
source attribution quality, and response consistency were assessed.
Results: GastroGPT-Deeplake achieved the highest overall performance score
(53.8±5.9; maximum 60 points), significantly outperforming GPT-4 (50.1±6.8,
p=0.028) and Claude (49.3±7.2, p=0.012). GastroGPT variants demonstrated marked
superiority over general-purpose models in hallucination rates (4.2–6.8% vs. 8.3–
12.1%, p=0.018) and source attribution quality (62.8–78.3% vs. 23.5–34.7%). In the
vector database comparison, Deeplake significantly outperformed Faiss and Chroma
(p<0.001). GastroGPT's mean scores across the five evaluation criteria ranged from 4.00 to 4.21, with the highest performance observed in guideline concordance
(4.21±0.63). Case complexity had a significant effect on diagnostic accuracy
(F=7.361, p=0.0008); diagnostic accuracy scores in high-complexity cases
(3.81±0.83) were significantly lower compared to low-complexity cases (4.43±0.70;
Cohen's d=0.803). No significant performance differences were observed across
disease prevalence categories or subspecialty areas. Inter-rater reliability was excellent
(ICC>0.99, Pearson r=0.814, Cronbach's α=0.802). ROC analysis yielded an AUC of
0.908 for the overall performance criterion.
Conclusion: GastroGPT-Deeplake demonstrated superiority over general-purpose
large language models in both overall performance and safety parameters through
domain-specific RAG integration. The advantages of RAG technology in reducing
hallucination risk and ensuring source traceability are of critical importance for the
safety of medical AI applications. Our findings support the preferential use of domainspecific approaches over general-purpose models in developing AI-assisted clinical
decision support systems for gastroenterology practice.
Keywords: Artificial intelligence, large language model, gastroenterology,
GastroGPT, retrieval-augmented generation, clinical decision support
system, hallucination, vector database