遇见数据集

S7 Fig -

收藏
Figshare2023-08-03 更新2026-04-28 收录
官方服务:

资源简介:

Deep learning (DL) techniques have seen tremendous interest in medical imaging, particularly in the use of convolutional neural networks (CNNs) for the development of automated diagnostic tools. The facility of its non-invasive acquisition makes retinal fundus imaging particularly amenable to such automated approaches. Recent work in the analysis of fundus images using CNNs relies on access to massive datasets for training and validation, composed of hundreds of thousands of images. However, data residency and data privacy restrictions stymie the applicability of this approach in medical settings where patient confidentiality is a mandate. Here, we showcase results for the performance of DL on small datasets to classify patient sex from fundus images—a trait thought not to be present or quantifiable in fundus images until recently. Specifically, we fine-tune a Resnet-152 model whose last layer has been modified to a fully-connected layer for binary classification. We carried out several experiments to assess performance in the small dataset context using one private (DOVS) and one public (ODIR) data source. Our models, developed using approximately 2500 fundus images, achieved test AUC scores of up to 0.72 (95% CI: [0.67, 0.77]). This corresponds to a mere 25% decrease in performance despite a nearly 1000-fold decrease in the dataset size compared to prior results in the literature. Our results show that binary classification, even with a hard task such as sex categorization from retinal fundus images, is possible with very small datasets. Our domain adaptation results show that models trained with one distribution of images may generalize well to an independent external source, as in the case of models trained on DOVS and tested on ODIR. Our results also show that eliminating poor quality images may hamper training of the CNN due to reducing the already small dataset size even further. Nevertheless, using high quality images may be an important factor as evidenced by superior generalizability of results in the domain adaptation experiments. Finally, our work shows that ensembling is an important tool in maximizing performance of deep CNNs in the context of small development datasets.

深度学习(Deep Learning, DL)技术在医学成像领域备受关注,尤其在利用卷积神经网络(Convolutional Neural Networks, CNNs)开发自动化诊断工具方面。视网膜眼底成像凭借其非侵入式采集的便捷性,格外适配这类自动化分析方案。当前基于CNNs的眼底图像分析研究,往往需要获取数十万级规模的海量数据集用于训练与验证。然而,数据驻留与数据隐私限制,阻碍了该方法在以患者隐私保护为硬性要求的医疗场景中的应用。 本研究展示了在小样本数据集场景下,利用深度学习模型从眼底图像中分类患者性别的实验结果——此前该特征被认为在眼底图像中不存在或无法被量化。具体而言,我们对残差网络-152(Residual Network-152, ResNet-152)模型进行微调,将其最后一层修改为适配二分类任务的全连接层。我们分别使用1个私有数据集(DOVS)与1个公开数据集(ODIR),开展多项实验以评估小样本场景下的模型性能。本研究基于约2500张眼底图像开发的模型,其测试集受试者工作特征曲线下面积(Area Under the Curve, AUC)得分最高可达0.72(95%置信区间:[0.67, 0.77])。相较于现有文献中的研究结果,尽管数据集规模缩减了近千倍,模型性能仅下降约25%。 本研究结果表明,即使是从视网膜眼底图像中进行性别分类这类颇具难度的二分类任务,也可依托极小规模的数据集实现。我们的域自适应(domain adaptation)实验结果显示,基于某一图像分布训练的模型,可良好泛化至独立的外部数据源,例如在DOVS数据集上训练、在ODIR数据集上测试的模型表现优异。此外,我们发现剔除低质量图像可能会进一步压缩本就有限的训练集规模,反而不利于CNN的训练。尽管如此,使用高质量图像仍是关键因素,域自适应实验中模型的泛化能力更优也印证了这一点。最后,本研究证实,集成学习(ensembling)是在小样本开发数据集场景下,最大化深度CNN模型性能的重要手段。

创建时间:
2023-08-03
二维码
社区交流群
二维码
科研交流群
商业服务