Multimodal Artificial Intelligence for Image-Based Prediction of Difficult Airway: A Prospective Observational Study
Multimodal Artificial Intelligence for Image-Based Prediction of Difficult Airway: A Prospective Observational Study
This prospective observational study will evaluate whether commonly available multimodal artificial intelligence models can predict difficult laryngoscopy and difficult intubation using standardized preoperative airway photographs. Adult patients scheduled for elective surgery requiring endotracheal intubation will undergo an eight-view preoperative airway photography protocol. The anonymized image sets will be assessed by ChatGPT, Gemini, and Grok using the same structured prompt. Their predictions will be compared with expert anesthesiologist image-based assessments, conventional airway evaluation findings, and prospectively recorded intraoperative airway outcomes. The primary aim is to determine the diagnostic performance of AI models for predicting difficult intubation. A key secondary aim is to evaluate their performance for predicting difficult laryngoscopy. The study is intended to explore whether image-based AI assessment may support preoperative airway risk stratification as a clinician-supervised screening tool.
Preoperative airway assessment is important for identifying patients at risk for difficult laryngoscopy or difficult intubation. However, conventional bedside airway predictors have limited accuracy when used alone. Multimodal artificial intelligence models may provide additional image-based information by evaluating visible anatomical features from standardized preoperative airway photographs.
In this prospective observational study, adult patients undergoing elective surgery requiring endotracheal intubation will be enrolled between June and September 2026. Each participant will undergo standardized eight-view airway photography during the pre-anesthetic evaluation. The image set will include frontal facial, lateral profile, maximal mouth opening, modified Mallampati, neck extension, and anterior neck views. Images will be anonymized before assessment.
The same image sets will be independently evaluated by multimodal AI models, including ChatGPT, Gemini, and Grok, using an identical structured prompt. The AI models will provide categorical and binary predictions for difficult laryngoscopy and difficult intubation based only on visible image-based anatomical features. No intraoperative outcome data, expert predictions, or conventional airway assessment results will be provided to the AI models.
AI-generated predictions will be compared with expert anesthesiologist image-based assessments, conventional airway evaluation parameters, and prospectively recorded intraoperative reference outcomes. Difficult laryngoscopy will be defined as Cormack-Lehane grade III or IV. Difficult intubation will be defined using objective intraoperative criteria, including more than one intubation attempt, need for bougie or stylet assistance, rescue use of video laryngoscopy or supraglottic airway device, intubation time exceeding 60 seconds, or Intubation Difficulty Scale score greater than 5.
The study will assess the sensitivity, specificity, positive predictive value, negative predictive value, accuracy, receiver operating characteristic performance, and agreement between AI models and expert anesthesiologist assessments. The findings may help clarify whether multimodal AI can serve as a clinician-supervised adjunct for preoperative difficult airway risk stratification.
Inclusion Criteria:
Exclusion Criteria: