Performance Analysis of Large Language Model Chatbots in Temporomandibular Disorders and Bruxism (Orthodontic and Oral Physiological Perspectives): Methodological Study Temporomandibular Bozukluklar ve Bruksizmde Büyük Dil Modeli Sohbet Robotlarının Performans Analizi (Ortodontik ve Oral Fizyolojik Perspektifler): Karşılaştırmalı Çalışma


BULUT A., AŞKIN M. B., Manav Ö. C.

Turkiye Klinikleri Journal of Medical Sciences, vol.46, no.2, pp.139-149, 2026 (Scopus, TRDizin)

  • Publication Type: Article / Article
  • Volume: 46 Issue: 2
  • Publication Date: 2026
  • Doi Number: 10.5336/medsci.2025-115242
  • Journal Name: Turkiye Klinikleri Journal of Medical Sciences
  • Journal Indexes: Scopus, Central & Eastern European Academic Source (CEEAS), CINAHL, EMBASE, TR DİZİN (ULAKBİM), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
  • Page Numbers: pp.139-149
  • Keywords: Chatbot, oral physiology, orthodontics, temporomandibular disorder
  • Yozgat Bozok University Affiliated: Yes

Abstract

Objective: This study aimed to evaluate and compare the informational reliability of 4 widely used large language model (LLM)-based chatbots ChatGPT-4, ChatGPT-3.5, Google Gemini, and Microsoft Copilot when responding to interdisciplinary questions concerning temporomandibular disorders (TMD) and bruxism from orthodontic and oral physiology perspectives. Material and Methods: A cross-sectional, comparative content analysis was conducted using 20 open-ended questions developed by orthodontic and physiology experts. Each chatbot generated 20 responses (80 in total), which were independently evaluated by blinded reviewers using a 4-domain rubric assessing scientific accuracy, depth and completeness, conceptual consistency, and referencing quality. Scores were analyzed with non-parametric statistical tests. Results: ChatGPT-4 achieved the highest overall performance (16.74±1.44), significantly outperforming ChatGPT-3.5 (14.52±1.74), Gemini (12.95±1.64), and Copilot (11.97±2.36) (p<0.001). “Post hoc” comparisons confirmed significant differences across most chatbot pairs, with ChatGPT-4 demonstrating superior accuracy, depth, and conceptual coherence. However, referencing quality was consistently low across all platforms. Conclusion: Advanced LLM-based chatbots, particularly ChatGPT-4, provide relatively accurate and coherent information on TMD and bruxism, though critical limitations persist in referencing and interdisciplinary integration. While these tools may support patient education, they should not substitute professional clinical expertise. Future improvements in domain-specific training and evidence-based reference integration are essential to enhance their reliability in dental and interdisciplinary practice.