Comparison of the Diagnostic Reliability of ChatGPT in Letournel-Judet Acetabulum Fracture Classification Based on Judet Radiographs According to Orthopaedic Residents
Comparison of the Diagnostic Reliability of ChatGPT in Letournel-Judet Acetabulum Fracture Classification Based on Judet Radiographs According to Orthopaedic Residents
This study aims to evaluate the diagnostic reliability of the multimodal artificial intelligence model ChatGPT-4o in classifying acetabular fractures using the Letournel-Judet classification system. The study retrospectively analyzed standardized pelvic radiographs (anteroposterior, iliac oblique, and obturator oblique) from 184 patients presenting with pelvic injuries. The diagnostic performance of ChatGPT-4o was compared against the independent assessments of two fourth-year orthopaedic residents and a reference standard established by an experienced trauma surgeon using multiplanar computed tomography (CT) and intraoperative findings. By utilizing a systematic radiographic checklist, the study assesses the AI (artificial intelligence) model's ability to identify key anatomical landmarks and integrate them into a final fracture pattern. This research aims to provide critical data on the current feasibility of using large language models as decision-support tools in complex orthopaedic trauma.
This retrospective observational study was conducted at Ankara Bilkent City Hospital to investigate the diagnostic accuracy of ChatGPT-4o in complex acetabular fracture classification. Imaging data included anonymized anteroposterior, iliac oblique, and obturator oblique radiographs from 184 patients.After approval from the Institutional Review Board, the investigators retrospectively analyzed patients presenting with pelvic injuries and undergoing surgical treatment at our center. Patients with pelvic injuries but with non-acetabular fractures, those with incomplete imaging, those with poor radiographic quality or those who had previously undergone pelvic surgery or trauma were excluded.
Assessment Protocol
Imaging from cases presenting with pelvic injuries and undergoing surgical treatment at our center was selected from the institution's radiology archive. Standard anteroposterior pelvic, iliac oblique and obturator oblique radiographs were collected for each case.
The article "Acetabular Fractures: Easier Classification with a Systematic Approach" by Brandser and Marsh was uploaded to Chat GPT. The prompt used when evaluating the cases was as follows: "In these AP pelvic, iliac oblique, and obturator oblique radiographs, identify the acetabulum fracture as defined by Judet, using the systematic approach in the provided article and answering each question on the radiological checklist individually." Each case was evaluated using the same prompt previously defined. Chat GPT was asked to answer each question individually based on the systematic approach defined in the article and ultimately indicate the type of fracture.These questions were designed to evaluate key radiographic landmarks and fracture components including:
Is there a fracture of the obturator ring? Is the ilioischial line disrupted? Is the iliopectineal line disrupted? Is there a fracture of the iliac wing? Is there a fracture of posterior wall? Does the fracture divide the acetabulum into top and bottom halves or front and back halves? Is there a spur sign?
Based on the responses to these questions, the final fracture type was determined according to the Letournel-Judet acetabular fracture classification system.
Cases included in the study were independently classified by two fourth-year orthopaedic residents. To ensure objectivity, assessments were based solely on anteroposterior (AP), obturator oblique, and iliac oblique radiographs, without access to CT scans or intraoperative findings. Also a trauma surgeon with over a decade of experience in pelvic trauma surgery independently classified cases. To establish the true fracture pattern, the experienced trauma surgeon serving as the reference standard had full access to multiplanar computed tomography (CT) scans and intraoperative records. In contrast, the AI model and the orthopaedic residents were completely blinded to any CT data, basing their evaluations solely on the plain Judet radiographs.
Statistical Analysis
Descriptive statistics were used to summarize patient demographics and fracture characteristics. Continuous variables (e.g., patient age) were expressed as mean ± standard deviation (SD) and range, whereas categorical variables (e.g., sex and fracture types) were presented as frequencies and percentages.
The primary outcome measure was the diagnostic accuracy of ChatGPT-4o, orthopaedic residents, and the trauma surgeon in classifying acetabular fractures according to the Letournel-Judet classification system. The reference standard was established using three-dimensional pelvic CT findings together with the final intraoperative assessment of the experienced trauma surgeon. Accuracy was calculated as the proportion of correctly classified fractures relative to the reference standard.
Interobserver agreement between evaluators and the reference standard was assessed using Cohen's kappa (κ) coefficient for pairwise comparisons. Kappa values were interpreted according to the criteria described by Landis and Koch: <0, poor agreement; 0.01-0.20, slight agreement; 0.21-0.40, fair agreement; 0.41-0.60, moderate agreement; 0.61-0.80, substantial agreement; and 0.81-1.00, almost perfect agreement.
Comparisons of diagnostic accuracy rates between ChatGPT-4o and human evaluators were performed using McNemar's test for paired categorical data. A p value of <0.05 was considered statistically significant. All statistical analyses were conducted using IBM SPSS Statistics version 29.0 (IBM Corp., Armonk, NY, USA).
Inclusion Criteria:
Exclusion Criteria: