Abstract
Purpose: To compare the accuracy of vertical cup-disc ratios (VCDR), ascertained by machine
learning (ML) versus human graders, from fundus images for glaucoma detection. This study
utilizes population-based data, with a disease prevalence and case-mix that is closer to a
real-world setting than conventional case-control studies, with the aim of developing
improved glaucoma screening tests.
Design: Cross-sectional analysis of a population-based study.
Participants: 6,304 participants of the EPIC-Norfolk Eye Study with color fundus images
gradable by humans and ML in both eyes.
Methods: VCDR was independently estimated from two-dimensional fundus images of
EPIC-Norfolk Eye Study participants by trained human graders (H-VCDR) and an externally
trained, open access, ML model (ML-VCDR). A neural network trained on 81,830
ophthalmologist-labeled images was used to generate pseudo-labels for over 100,000 UK
Biobank images, on which ML-VCDR was subsequently trained. Glaucoma status was
ascertained by tertiary center specialist examination. Predictive performance of VCDR for
glaucoma status was examined using logistic regression. ML-VCDR estimates were
additionally compared to a popular open-source ML model (AutoMorph) and scanning laser
ophthalmoscopy (Heidelberg Retinal Tomography (HRT)).
Main Outcome Measures: Area Under the Receiver Operated Characteristic Curve (AUROC),
explained variance (McFadden’s pseudo-R2).
Results: Of 6,304 participants (mean age 68 years; 57% women), 696 had glaucoma or
suspect status in at least one eye. For left eyes, H-VCDR and ML-VCDR explained 17% (95% CI
14.7 - 20.4) and 31% (95% CI 27.9 - 33.9) of glaucoma status variance and had an area under
the ROC curve (AUROC) of 79% (95% CI 76.8 - 81.2) and 88% (95% CI 86.3 - 89.1),
respectively. Right eye H-VCDR and ML-VCDR explained 20% (95% CI 16.9 - 22.6) and 35%
(95% CI 32.4 - 37.9) of the variance, and had an AUROC of 81% (95% CI 78.6 - 82.5) and 90%
(95% CI 88.7 - 91.0), respectively. ML-VCDR also performed better than AutoMorph and HRT
at predicting glaucoma status from VCDR estimations.
Conclusions: In this population-based setting, ML far outperformed trained human graders
at predicting specialist-ascertained glaucoma status from fundus images. This provides
promise for ML-supported strategies for glaucoma population screening.