Figures Abstract Large domain-specific foundation models have been widely adopted for retinal image analysis, yet systematic evidence for their advantage over compact general-purpose architectures remains scarce. We benchmarked nine model configurations spanning 22.8M to 303M parameters (vision transformers, hierarchical Swin Transformers, ConvNeXt, and the domain-specific RETFound models) across four tasks: OCT 8-class disease classification, and three fundus photography tasks (DME severity, glaucoma detection, and DR severity grading). All models were evaluated under identical training conditions, with both pretrained (on natural-domain image datasets) and from-scratch initializations compared using Mann-Whitney U tests. Pretraining improved accuracy by 5.18–18.41 percentage points across all tasks (p < 0.05 throughout), with larger benefits for CFP modalities and harder tasks. Compact hierarchical models (27–29M parameters) matched or exceeded larger architectures on three of four tasks. For instance, the SwinV2-tiny architecture ranked first on OCT, DME, and GL classification. The domain-specific RETFound model (303M) achieved the highest accuracy only on the most challenging task (DR severity grading, where the most severe class is underrepresented at 8% of images), where it outperformed the best compact model by 1.54 percentage points. These results indicate that compact general-purpose models may be sufficient for most retinal classification benchmarks, and that domain-specific foundation models may add higher value mainly for severity grading tasks with skewed class distributions. Citation: Isztl D, Spitznagel T, Somfai GM, Santos R (2026) Compact vision models match domain-specific foundation models for several retinal imaging classification tasks: A systematic benchmark. PLoS One 21(8): e0356202. https://doi.org/10.1371/journal.pone.0356202 Editor: Jia-Lang Xu, National Taichung University of Science and Technology, TAIWAN Received: April 14, 2026; Accepted: July 29, 2026; Published: August 18, 2026 Copyright: © 2026 Isztl et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium,
Compact vision models match domain-specific foundation models for several retinal imaging ...
Read the original article
journals.plos.org →