Prospective Multi-Center Evaluation of AI-Assisted Interpretation of Ultra-Widefield Retinal Images in a Multi-Reader Crossover Study
Prospective Multi-Center Evaluation of AI-Assisted Interpretation of Ultra-Widefield Retinal Images in a Multi-Reader Crossover Study
The goal of this prospective observational study is to evaluate the impact of artificial intelligence (AI) assistance on clinician interpretation of ultra-widefield (UWF) retinal images.
The main questions it aims to answer are:
whether AI assistance improves the diagnostic performance of ophthalmologists in detecting retinal findings on UWF retinal images; whether AI assistance improves sensitivity, specificity, and inter-reader agreement across clinicians with different levels of experience.
Approximately 600 UWF retinal images prospectively collected from multiple ophthalmic centers in China will be included. Images will be independently annotated by expert retinal specialists to establish reference labels for retinal finding categories.
Four ophthalmologists with different levels of clinical experience, including one senior retinal specialist and three junior ophthalmologists, will participate in a crossover multi-reader study.
For each clinician, the dataset will be randomly divided into two equal subsets. During the first reading session, clinicians will evaluate one subset without AI assistance and the other subset with AI assistance. After a washout interval of at least two weeks, the reading conditions will be reversed in a second reading session with independently randomized image order.
Under the AI-assisted condition, clinicians will be provided with category-level AI prediction probabilities for retinal findings. No localization maps, heatmaps, segmentation overlays, or automated diagnostic recommendations will be displayed. Clinicians will retain full autonomy over final decisions.
Reader performance under AI-assisted and unaided conditions will be compared using expert reference annotations as the ground truth.
This study is a prospective multi-center observational reader study designed to evaluate the impact of artificial intelligence (AI) assistance on clinician interpretation of ultra-widefield (UWF) retinal images.
Approximately 600 UWF retinal images will be prospectively collected from multiple ophthalmic centers in China. Images will be acquired using clinically routine UWF retinal imaging systems and will include a broad spectrum of retinal diseases and retinal findings encountered in real-world clinical practice.
All images will undergo independent expert annotation by retinal specialists to establish reference labels for retinal finding categories. These expert annotations will serve as the reference standard for subsequent performance evaluation.
Four ophthalmologists with different levels of clinical experience will participate in the reader study, including:
one senior retinal specialist with approximately five years of retinal clinical experience; three junior ophthalmologists with approximately two years of ophthalmology residency training.
A randomized crossover multi-reader design will be implemented to minimize recall bias and balance reading conditions.
For each clinician, the image dataset will be randomly divided into two equal subsets (subset A and subset B; approximately 300 images each).
During Round 1:
subset A will be interpreted without AI assistance; subset B will be interpreted with AI assistance.
After a washout interval of at least two weeks, the reading conditions will be reversed during Round 2:
subset A will be interpreted with AI assistance; subset B will be interpreted without AI assistance.
Image order will be independently randomized for each session and each clinician.
Under the unaided condition, clinicians will evaluate retinal images using standard clinical interpretation without AI output.
Under the AI-assisted condition, clinicians will receive category-level AI prediction probabilities for retinal finding categories. The AI output will provide probabilistic confidence scores only and will not include lesion localization maps, heatmaps, segmentation overlays, or automated binary recommendations.
Clinicians will remain blinded to the expert reference labels and to the interpretations of other readers. Final diagnostic decisions will be independently determined by each clinician.
The primary analysis will compare diagnostic performance between unaided and AI-assisted conditions, including sensitivity, specificity, area under the receiver operating characteristic curve (AUC), and inter-reader agreement.
Inclusion Criteria:
Exclusion Criteria: