This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment. The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs. Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.
Inclusion Criteria:
Exclusion Criteria:
mariam.ahmed.hosam@dentistry.cu.edu.eg00201110913251
Diagnostic Accuracy of Educated Large Language Models in Endodontic Diagnosis and Case Difficulty Assessment
LLM Performance in Endodontic Diagnostics
Diagnostic Accuracy of Oral Images, OPGs, and Questionnaires vs. Clinical Assessment for Periodontal Disease
Diagnostic Accuracy of Oral Images, OPGs, Biomarkers and Questionnaires vs. Clinical Assessment for Periodontal Disease (PostNCT07164573)
Development and Diagnostic Accuracy of a Deep Learning Model for Root Canal Curvature Analysis in Mandibular Molars Using CBCT Scans: A Diagnostic Accuracy Study
Development and Validation of a Deep Learning Model to Predict Endodontic Retreatment Difficulty From Periapical Radiographs
Evaluation of AI Large Models for Diagnosis and Treatment in Real-World Cases: Multicenter Retrospective Study
Ophthalmic Diseases and AI: an RCT Study