DeepAdapter: a generalisable algorithm integrating self-supervised learning and unsupervised domain adaptation for robust retinopathy of prematurity screening

  1. Jiaman Zhao1,2,
  2. http://orcid.org/0009-0009-7533-5642Longhui Li1,
  3. Zhenzhe Lin1,
  4. Mingjie Luo1,
  5. Dongyuan Yun1,2,
  6. Jingyi Wen1,
  7. Lixue Liu1,
  8. Jianqiao Li3,
  9. Fabao Xu3,
  10. Suzhen Xie4,
  11. Meirong Wei5,
  12. Baohai Liu6,
  13. http://orcid.org/0000-0002-4853-2474Zhenzhen Liu1,
  14. Duoru Lin1,
  15. http://orcid.org/0000-0003-4672-9721Haotian Lin1,7
  1. 1Sun Yat-sen University, Zhongshan Ophthalmic Center, State Key Laboratory of Ophthalmology, Guangzhou, Guangdong, China
  2. 2School of Biomedical Engineering, Sun Yat-sen University, Shenzhen, China
  3. 3Department of Ophthalmology, Qilu Hospital, Shandong University, Jinan, China
  4. 4Guangdong Women and Children Hospital, Guangzhou, China
  5. 5Department of Ophthalmology, Maternal and Children’s Hospital, Liuzhou, China
  6. 6Department of Ophthalmology, Maternal and Children’s Hospital, Linyi, China
  7. 7Hainan Eye Hospital and Key Laboratory of Ophthalmology, Zhongshan Ophthalmic Center, Sun Yat-sen University, Haikou, China
  1. Correspondence to Professor Haotian Lin; gddlht{at}aliyun.com; Professor Duoru Lin; lindr5{at}mail.sysu.edu.cn; Dr Zhenzhen Liu; liuzhenzhen{at}gzzoc.com

Abstract

Aims To develop and validate DeepAdapter, a novel deep learning algorithm that integrates self-supervised learning (SSL) and unsupervised domain adaptation (UDA) to enhance model generalisability for retinopathy of prematurity (ROP) screening.

Methods DeepAdapter was developed in two stages. First, SSL was applied to 500 000 unlabelled infantile colour fundus photographs (CFPs) retrospectively collected from four Chinese clinical centres to learn general fundus representations. Second, a much smaller labelled dataset of Normal and ROP CFPs was constructed, and a UDA module was incorporated to mitigate domain shift, defined as distribution differences between training and deployment datasets that may impair model performance. Correlation Alignment (CORAL) distance was employed to quantify these domain shifts. External validation was performed using data from an independent clinical centre and two public multiethnic datasets.

Results A total of 500 000 non-overlapping unlabelled CFPs and 39 456 labelled CFPs were included. Compared with the supervised method, DeepAdapter decreased CORAL distance from 0.297 to 3.550×10−6. Both algorithms achieved high internal accuracy (0.989) in ROP prediction, but DeepAdapter significantly outperformed the supervised method on external testing (accuracy: 0.828 (95% CI 0.823 to 0.833) vs 0.739 (95% CI 0.733 to 0.745), p <0.001) and on the public cross-ethnicity dataset (accuracy: 0.839 (95% CI 0.818 to 0.859) vs 0.813 (95% CI 0.791 to 0.835), p=0.014). A web-based system integrating quality control and ROP diagnosis was developed for large-scale, multicentre clinical screening.

Conclusions DeepAdapter effectively mitigated domain shift and significantly improved model generalisability for ROP screening, providing a valuable reference for developing generalisable models in other medical specialties.

  • Retinopathy of Prematurity
  • Imaging
  • Artificial Intelligence

Data availability statement

Data are available upon reasonable request. Requests to access research materials and code for research purposes should be directed to the corresponding author. A demonstration video of the web-based system developed in this study is provided in the online supplementary materials, and a demo version is available at https://deepadapter.cn.

https://creativecommons.org/licenses/by-nc/4.0/

This is an open access article distributed in accordance with the Creative Commons Attribution Non Commercial (CC BY-NC 4.0) license, which permits others to distribute, remix, adapt, build upon this work non-commercially, and license their derivative works on different terms, provided the original work is properly cited, appropriate credit is given, any changes made indicated, and the use is non-commercial. See: https://creativecommons.org/licenses/by-nc/4.0/.

Statistics from Altmetric.com

Request Permissions

If you wish to reuse any or all of this article please use the link below which will take you to the Copyright Clearance Center’s RightsLink service. You will be able to get a quick price and instant permission to reuse the content in many different ways.

  • Retinopathy of Prematurity
  • Imaging
  • Artificial Intelligence

WHAT IS ALREADY KNOWN ON THIS TOPIC

  • Generalisation remains a major barrier for medical artificial intelligence (AI) applications, as models achieving expert-level performance on internal datasets often suffer substantial performance degradation on unseen data due to domain shift. Although AI-based retinopathy of prematurity (ROP) screening models have been developed, few studies have addressed the generalisation issues caused by domain shift.

WHAT THIS STUDY ADDS

  • By quantitative analysis of domain shift in infantile colour fundus photographs using Correlation Alignment distance, we observed notable discrepancies across clinical centres. Compared with the supervised method, DeepAdapter substantially mitigated domain shift and achieved significantly better performance on external centres and public datasets for ROP detection.

HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY

  • DeepAdapter effectively mitigates domain shift and demonstrates superior generalisability for ROP detection in diverse settings. Given the ubiquity of domain shift in medical AI, DeepAdapter may serve as a valuable reference for developing generalisable models across other medical specialties.

Introduction

Retinopathy of prematurity (ROP) is the most common and preventable cause of infantile blindness in developing countries.1 Early detection and intervention are essential to prevent irreversible vision loss, yet nearly 40% of children worldwide lack timely treatment.2 Globally, approximately 13.4 million newborns are born preterm each year,3 with an ROP prevalence of 31.9%.4 However, the shortage and uneven distribution of paediatric ophthalmologists hinder large-scale ROP screening efforts.5 Therefore, finding a feasible solution for ROP screening under limited medical resources remains a pressing public health priority.

In recent years, deep learning (DL) algorithms have enabled end-to-end ROP detection,6–8 significantly improving the efficiency of fundus image analysis and diagnosis. Most of these studies have adopted supervised methods and achieved high performance on specific datasets, but have paid little attention to improving generalisability across clinical centres and ethnic groups. Such limited generalisability is typically driven by domain shift—the discrepancy between development datasets and real-world deployment data, which frequently precipitates performance degradation.9 For instance, a previous study developed a DL model using 7414 colour fundus photographs (CFPs) from Homerton University Hospital in London, achieving an internal validation area under the curve (AUC) of 0.986. However, when externally validated on an American dataset comprising 100 CFPs, the AUC dropped to 0.808.10 This variability limits the clinical application of artificial intelligence (AI) in ROP screening by constraining generalisability.

Supervised deep neural networks are highly sensitive to patterns in training data, including confounding factors such as pixel-level acquisition variability and differences in fundus pigmentation,11 resulting in performance deterioration on unseen data. In paediatric fundus screening, domain shifts arise from substantial heterogeneity across centres, including differences in fundus camera brands, population and environmental characteristics such as fundus pigmentation, birth weight and gestational age, as well as operator-dependent variations in image acquisition. In addition, our prior study found that 49.4% of CFPs exhibited quality defects due to poor infant compliance,12 further limiting model generalisability. These challenges highlight the need for more generalisable and adaptable AI solutions.

Recent advances in self-supervised learning (SSL) and unsupervised domain adaptation (UDA) offer promising solutions to the limited generalisability of conventional supervised methods. SSL leverages large scale of unlabelled data through pretext tasks, enabling models to learn general representations directly from the data.13 UDA, a branch of transfer learning, mitigates domain shifts between source and target domains by aligning shared structures or latent representations.14 Improving real-world clinical generalisability requires addressing two complementary challenges: enhancing the robustness of disease-relevant feature representations and reducing non-disease-related confounding factors that drive domain shift. These objectives align with the respective strengths of SSL and UDA. Although SSL or UDA have been studied in medical AI,15–17 their synergistic integration remains underexplored. To address this gap, we propose DeepAdapter, an algorithm that integrates SSL and UDA to harness their complementary strengths for developing a generalisable AI model for ROP screening.

Materials and methods

Study design and participants

Figure 1 illustrates the overall research process of this study, which is divided into four stages: image acquisition and processing, dataset construction, model development, and model evaluation. To develop and test an ROP screening model employing the DeepAdapter algorithm, we retrospectively collected unlabelled infantile CFPs from newborn fundus screening cohorts at four clinical centres: Guangdong Women and Children Hospital, Liuzhou Maternity and Child Healthcare Hospital, Linyi Maternity and Child Healthcare Hospital, and Zhongshan Ophthalmic Centre (ZOC). These CFPs were captured using Retcam (Natus Medical, Pleasanton, USA), PanoCam (Visunex Medical Systems, San Ramon, USA) or Orthocone (RS-B002, China). Data from Guangdong Women and Children Hospital were retained as the external-centre test dataset, and data from the remaining centres were randomly divided into training and validation datasets.

Figure 1

Overview of the study workflow for developing and evaluating the DeepAdapter model for retinopathy of prematurity screening. The pipeline comprises four main stages: (1) image acquisition and quality control of CFPs, (2) dataset construction with expert annotation, (3) model development integrating self-supervised learning and unsupervised domain adaptation, and (4) internal and external evaluation. CFPs, colour fundus photographs.

Data preparation

All CFPs were processed using a DL-based image quality assessment and enhancement system (DeepQuality) previously developed by ZOC.18 Images with a quality score below 0.2 were excluded, resulting in a retention rate of 89.8%. During model training, data augmentation was applied to increase data volume and diversity and to enhance model robustness, including random brightness and contrast adjustments, gamma correction, addition of Gaussian noise, random rotations, and horizontal or vertical flips. As ROP lesions typically develop in the peripheral retina, particularly at the junction between vascularised and avascular zones, uniform cropping, a common preprocessing step for many convolutional neural networks to ensure fixed-size inputs, may inadvertently exclude these regions. To address this, we employed a resizing and padding strategy to standardise image dimensions while preserving lesion information.

For downstream training and validation of DeepAdapter, we constructed a dataset of CFPs labelled as ‘Normal’ or ‘ROP’ according to the International Classification of Retinopathy of Prematurity guidelines,19 which define ROP by two key features: (1) basic lesions, characterised by abnormal vascular proliferation at the retinal vascular–avascular junction and classified into five stages; (2) plus disease, involving pathological dilation and tortuosity of posterior retinal vessels, including both preplus and plus disease. Images exhibiting any stage of basic lesions and/or plus/preplus disease were labelled ‘ROP’, whereas images devoid of these features were labelled ‘Normal’. All CFPs were annotated and re-examined by three ophthalmologists from the ZOC, Sun Yat-sen University. In cases of disagreement, a senior expert with over 15 years of clinical experience made the final decision.

Development and evaluation of DeepAdapter

The development pipeline for DeepAdapter and the conventional supervised method is illustrated in online supplemental figure S1. Both methods are based on the ResNet50 architecture, selected for its proven performance and computational efficiency, making it well-suited for scalable clinical deployment. To enhance the model’s ability to capture disease-relevant features, an attention mechanism was inserted between the last two residual blocks of ResNet50, with Grad-CAM (online supplemental figure S2) highlighting clinically relevant ROP regions such as the retinal ridge and tortuous vessels. This mechanism generates attention maps by computing attention scores, enabling the model to focus dynamically on salient regions and to fuse attention features with the original representation. Additionally, a bottleneck layer preceding the classification head was incorporated exclusively during the UDA phase.

Supplemental material

DeepAdapter began with pretraining on large-scale unlabelled CFPs using Momentum Contrast (MoCo) V.2,20 an efficient and scalable SSL algorithm designed to extract general fundus representations. During data augmentation, we designed a scale-and-padding strategy tailored for infantile scenarios, instead of the commonly used crop-centre augmentation, to standardise image dimensions while preserving peripheral lesion information, as critical diagnostic features for ROP often appear in the retinal periphery. Then the model was fine-tuned to acquire disease-specific classification capabilities. To address domain shift when applied to external clinical centres (target domains), DeepAdapter further incorporated a UDA algorithm—the Deep Subdomain Adaptation Network (DSAN)21 during fine-tuning.

Specifically, MoCo V.220 constructs proxy tasks based on positive pairs (different augmented views of the same image) and negative pairs (unrelated samples). The objective is to minimise the distance between positive pairs while maximising the distance between negative pairs in the encoded feature space. This encourages the model to learn intrinsic and generalisable features from the data itself, rather than label-driven representations. For domain adaptation, DSAN21 aligns class-wise feature distributions between the source (training) and target (external) domains by minimising the local maximum mean discrepancy (LMMD), without requiring labels from the target domain. During this phase, high-level encoded features pass through a bottleneck layer to obtain compact feature representations, which are then used to compute the transfer regularisation term based on LMMD, a soft-label mechanism that probabilistically aligns categories across domains. The overall training objective during domain adaptation fine-tuning jointly optimises the classification loss and the LMMD loss, as detailed in online supplemental appendix 1.

In contrast, the supervised method was pretrained on large-scale labelled datasets and fine-tuned on the same labelled source-domain CFPs, serving as a control to evaluate the impact of integrating SSL and UDA in DeepAdapter. Both models were trained and tested using quality-assessed and enhanced infantile CFPs and ultimately output diagnostic predictions. The final model was selected based on its performance on the validation set, with accuracy predefined as the primary selection criterion. In addition, we compared DeepAdapter with the state-of-the-art (SOTA) foundation model RETFound.22 Specifically, we initialised RETFound with its official pretrained weights (ViT-Large backbone) and fine-tuned it on the same internal training dataset for a fair comparison. All experiments were conducted using Python (V.3.8) and the PyTorch framework on Ubuntu V.20.04.6 LTS with four NVIDIA Tesla V100 GPUs (Graphics Processing Units).

Evaluation of domain shift

To quantify domain shift, we employed the Correlation Alignment (CORAL) distance,23 which measures the distributional discrepancy between source and target feature representations, with calculation details provided in the online supplemental appendix 2. Higher values represent larger discrepancies. To further visualise cross-domain feature alignment, we employed t-distributed stochastic neighbour embedding (t-SNE), a non-linear dimensionality reduction technique that projects high-dimensional features into two dimensions while approximately preserving the relative distances between nearby points, enabling intuitive visualisation of the feature distributions. Greater overlap between training and target domain samples indicates better alignment.

Statistical analysis

To comprehensively evaluate and compare the predictive performance of DeepAdapter with that of the supervised method, we employed metrics including sensitivity, specificity and F1-score, the area under the receiver operating characteristic (ROC) curve (AUC), and the area under the precision-recall (PRC) curve (AUPRC). AUC and AUPRC summarise overall discriminative performance in a threshold-independent manner. For metrics requiring a predefined threshold (eg, accuracy, sensitivity and specificity), a fixed threshold of 0.5 was applied across all experiments to ensure fair and consistent evaluation. Non-parametric bootstrap hypothesis testing was conducted to assess the statistical significance of performance differences, with p<0.05 considered statistically significant.

Results

Characteristics of the datasets

A total of 500 000 unlabelled CFPs were collected for the SSL stage of DeepAdapter development. In addition, 38 502 CFPs were annotated for training, validation and testing of both DeepAdapter and the supervised method. All CFPs were anonymised and assessed solely based on the images’ content, without involving any patient-identifiable information. The internal datasets were randomly split into training and validation sets in a 7:3 ratio. Additionally, 954 CFPs were curated from public datasets24 25 according to the same criteria as the training dataset to further evaluate model performance. The supervised method was trained and evaluated on the same annotated dataset to ensure a fair comparison. Details on annotation types and disease distributions across datasets are summarised in online supplemental table S1. Information about imaging devices used at each centre, along with demographic data including average gestational age and birth weight, is summarised in table 1. Infants with ROP had lower average birth weight and gestational age compared with those in the Normal group.

Table 1

Fundus imaging systems and infant cohort demographics across labelled datasets from multiple clinical centres

Feature distribution discrepancies across datasets

To quantify the domain shift of CFPs encoded by DeepAdapter and the supervised method across datasets, we computed CORAL distance based on second-order statistics (online supplemental table S2). For the supervised method, the value between the training and internal validation sets was minimal (1.335×10−5), indicating that UDA was unnecessary in this setting. In contrast, the value between the training and external test sets was substantially larger for the supervised method (0.297), indicating notable domain shift, which was likely resulting from variations in imaging procedures and patient demographics. However, DeepAdapter markedly reduced it by incorporating UDA (3.550×10−6). A consistent conclusion was observed for the public dataset (from 0.145 to 3.336×10−6). These results highlight substantial heterogeneity of infantile CFPs across clinical centres and demonstrate that DeepAdapter effectively mitigates such shifts, thereby enhancing model generalisability across populations and institutions.

Moreover, t-SNE visualisation (figure 2) revealed that DeepAdapter achieves better domain alignment than the supervised method, with greater overlap between training and target domain samples, indicating improved generalisation across domains. Specifically, t-SNE plots of the training and public datasets indicate that DeepAdapter achieves better alignment between samples from the two sources in both the Normal and ROP classes compared with the supervised method.

Figure 2

t-SNE visualisations of feature distributions across domains from the supervised method (A) and DeepAdapter (B). Each point represents a sample, coloured by class label (ROP or Normal) and shaped by dataset origin (train vs validation/test/public). Within the same class (colour), a greater overlap between samples from different domains (shapes) indicates better feature similarity between the source and target data. ROP, retinopathy of prematurity; t-SNE, t-distributed stochastic neighbour embedding.

Model performance evaluation

In the internal validation experiment, DeepAdapter and the supervised method were evaluated on a validation set drawn from the same distribution as the training set, with quantitative analysis confirming negligible domain shift (CORAL distance=1.335×10−5). As shown in table 2 and online supplemental figure S3), both methods achieved excellent performance on ROP prediction, with an accuracy of 0.99 and near-perfect ROC curves. Non-parametric bootstrap resampling showed no statistically significant performance differences across metrics (p>0.05).

Table 2

Performance of DeepAdapter and the supervised method in retinopathy of prematurity detection across internal, external and public datasets

In contrast, the external test set, collected from a different clinical centre in China, exhibited a significant domain shift from the training data, as confirmed by quantitative distribution analysis. Accordingly, DeepAdapter was fine-tuned with DSAN for this setting. As shown in table 2 and figure 3A, DeepAdapter significantly outperformed the supervised method across all metrics (accuracy: 0.828 vs 0.739; sensitivity: 0.882 vs 0.837; F1-score: 0.777 vs 0.701; AUC: 0.976 vs 0.942; specificity: 0.923 vs 0.884; all p<0.05).

Figure 3

Performance comparison between DeepAdapter and supervised method on external and public test datasets for retinopathy of prematurity prediction. (A) External test set from a different clinical centre in China. (B) Public multiethnic test set (South Asian and Caucasian populations). Left panels: ROC curves. Right panels: bar plots of evaluation metrics with 95% CIs. ***p<0.001; **p<0.01; *p<0.05; ns: not significant (p>0.05). ACC, accuracy; AUC, area under the curve; AUPRC, area under the precision-recall curve; ROC, receiver operating characteristic.

The public dataset was obtained from two open-access repositories representing distinct ethnic populations: South Asian (Indian) and Caucasian (Czech Republic). These populations differ genetically and phenotypically from the predominantly East Asian training set from China, with quantitative analysis confirming a substantial domain shift. DeepAdapter remained superior to the competing methods, as shown in table 2 and figure 3B (accuracy: 0.839 vs 0.813, p=0.014; sensitivity: 0.790 vs 0.748, p=0.002; F1-score: 0.783 vs 0.713, p<0.001; AUC: 0.934 vs 0.932, p=0.309; specificity: 0.902 vs 0.899, p=0.319). While specificity was comparable for identifying non-ROP cases, DeepAdapter showed higher sensitivity for positive cases, highlighting its clinical value for risk management.

In addition, comparisons with the SOTA model RETFound demonstrate that DeepAdapter significantly outperforms RETFound (online supplemental table S3). Furthermore, we replaced the UDA component of DeepAdapter with several alternative methods, and the results confirm the superiority of DSAN in our framework (online supplemental table S4).

Deployment and clinical integration of DeepAdapter

To deploy DeepAdapter in a new centre, a one-time, offline unsupervised alignment should be performed using a small batch of unlabelled data from the target domain to achieve optimal local performance. This process requires minimal additional training, no extra annotations and modest computational resources (see online supplemental appendix 3 for details). We have also developed an integrated system that combines infantile fundus image quality assessment, image enhancement and ROP prediction. After the images are uploaded, the system automatically evaluates quality, enhances the images, reassesses quality and generates ROP predictions.

Discussion

In this study, we present DeepAdapter, a unified DL algorithm that integrates SSL and UDA to enhance model generalisation and support clinical deployment of AI for ROP screening. Quantitative analysis of domain shift based on CORAL distance demonstrated substantial heterogeneity between the training data and external datasets. DeepAdapter effectively mitigates domain shifts by joint optimisation of classification loss and the LMMD loss, and consistently and significantly outperforms the conventional supervised method in detecting ROP across external and public datasets. These results highlight DeepAdapter’s potential to improve the reliability and generalisability of medical AI in real-world clinical settings.

Previous studies have mostly focused on improving model performance within specific datasets through supervised methods,6–8 while paying little attention to enhancing generalisability in practical applications. For example, Bai et al developed a DL algorithm (ROP.AI) using a single Australian cohort, achieving an accuracy of 97.3%, AUC of 0.993 and sensitivities and specificities of 96.6% and 98.0%, respectively.26 However, in a subsequent multicentre validation study conducted in Australia with 8052 images, the model’s performance dropped substantially, with an AUC of 0.75 and sensitivity and specificity of 84% and 43% after threshold optimisation.27 Compared with previous studies, we propose a novel algorithm, DeepAdapter, which integrates SSL and UDA to enhance the generalisability of ROP screening across diverse clinical settings. Unlike conventional supervised approaches that rely on label-associated features, SSL enables DeepAdapter to learn general representations from large-scale unlabelled data, while UDA mitigates domain shift to further improve generalisability. Moreover, we conducted both quantitative and interpretative analyses of domain shift across multicentre, multinational datasets. These analyses revealed substantial data heterogeneity and demonstrated DeepAdapter’s effectiveness in mitigating this challenge through UDA integration.

The performance improvement of DeepAdapter is probably attributable to the complementary strengths of SSL and UDA in addressing two core challenges of domain generalisation: representation robustness and domain shift (see online supplemental appendix 4 for theoretical details). SSL leverages large-scale unlabeled data to learn domain-invariant characteristics, such as geometric patterns and local relationships. In parallel, UDA minimises the domain shift between the source and target domains, thereby reducing the influence of non-disease-related confounding factors. While SSL establishes a robust and generalisable feature backbone, UDA fine-tunes the decision boundaries to better fit the target domain without requiring labelled target data. This integration effectively addresses the primary causes of cross-domain performance degradation: overfitting to source-specific patterns, misalignment of feature distributions and unstable decision boundaries in ambiguous regions.

Despite the promising results, several limitations and challenges warrant further investigation. First, although DeepAdapter was retrospectively evaluated across clinical centres and public datasets, its real-world effectiveness requires prospective validation. Second, while DeepAdapter demonstrated feasibility and superiority in ROP, the leading cause of blindness in infants, further validation is required for a broader disease spectrum. Nevertheless, it could theoretically be applied to other infantile retinal diseases. In addition, this study focuses on the screening stage, and future work will explore models for ROP severity grading to further support clinical decision-making.

In conclusion, this study introduced DeepAdapter, a generalised algorithm that effectively mitigates domain shifts and demonstrates strong generalisability in external clinical settings for ROP detection. Given the ubiquity of domain shift issues in medical AI, DeepAdapter offers a transferable paradigm for developing generalisable models, holding the potential to accelerate the clinical translation of AI across diverse medical specialties.

Data availability statement

Data are available upon reasonable request. Requests to access research materials and code for research purposes should be directed to the corresponding author. A demonstration video of the web-based system developed in this study is provided in the online supplementary materials, and a demo version is available at https://deepadapter.cn.

Ethics statements

Patient consent for publication

Not applicable.

Ethics approval

This study adhered to the principles of the Declaration of Helsinki and was approved by the Institutional Review Board of Zhongshan Ophthalmic Center at Sun Yat-sen University (IRB-ZOC-SYSU, identifier: 2023KYPJ029).

Read the full text or download the PDF:

Log in using your username and password