Multi-Label Retinal Diseases (MuReD) Dataset
收藏资源简介:
Abstract: Early detection of retinal diseases is one of the most important means of preventing partial or permanent blindness in patients. One of the major stumbling blocks for manual retinal examination is the lack of a sufficient number of qualified medical personnel per capita to diagnose diseases. Computer-aided diagnosis systems (CAD) have proven to be very effective in helping physicians reduce the time taken to make a diagnosis and minimize variability in image interpretation. Still, they are not flexible enough to accommodate the simultaneous presence of multiple retinal diseases, which is a common situation in real-world applications. In the past years, few datasets that focus on the classification of numerous retinal pathologies present at the same time, i.e., multi-label classification have been proposed, but there are some shared problems with all of them, such as a narrow range of pathologies to classify, high level of class imbalance, low amount of samples for the underrepresented labels, no assurance in image quality, among others. All these problems hinder the performance of any model trained with these datasets, which leads to poor robustness, lack of generalization, and reduced trustability in its predictions. To address these problems, we constructed the Multi-Label Retinal Diseases (MuReD) dataset, using images collected from three different state-of-the-art sources, i.e., ARIA, STARE, and RFMiD datasets, and performing a sequence of post-processing steps to ensure the quality of the images, a wide range of diseases to classify, and a sufficient number of samples per disease label. The MuReD dataset consists of 2208 images with 20 different labels, with varying image quality and resolution. At the same time, ensuring a minimal degree of quality in the data, with a sufficient number of samples per label. To the best of our knowledge, the MuReD dataset, is the only publicly available dataset that applies a sequence of post-processing steps to ensure the quality of the images, the variety of pathologies, and the number of samples per label, resulting in increased data quality and a significant reduction of the class imbalance present in the publicly available datasets. It is envisaged that the MuReD dataset will enable the creation of more robust, general, and trustable models for the automatic detection and classification of retinal diseases. Files Description: 1. The file "train_data.csv" contains the images from the training set, along with the 20 different labels. 2. The file "val_data.csv" contains the images from the validation set, along with the 20 different labels. 3. The folder "images" contains all the images that compose the MuReD dataset. The images come in two different formats, i.e., .tiff and .png. There is no single image resolution. Given that the images come from different sources, resolution can vary from 520x520 to 3400x2800 depending on the source of the image.
摘要:早期检出视网膜疾病是预防患者部分或永久性失明的最重要手段之一。人工视网膜检查的主要障碍之一是人均合格医护人员数量不足,难以开展疾病诊断。计算机辅助诊断系统(CAD)已被证实可有效帮助医师缩短诊断时长、降低图像判读的变异性,但此类系统灵活性不足,无法适配同时存在多种视网膜疾病的临床常见场景。近年来,虽有少量面向多病理共存场景(即多标签分类)的视网膜疾病分类数据集被提出,但此类数据集普遍存在若干共性问题:待分类病理种类范围狭窄、类别不平衡程度较高、稀有标签对应样本量不足、无法保证图像质量等。上述问题均会制约基于此类数据集训练的模型性能,导致模型鲁棒性不足、泛化能力缺失且预测可信度下降。 为解决上述问题,本研究构建了多标签视网膜疾病(MuReD)数据集。该数据集的图像源自ARIA、STARE与RFMiD三类当前主流公开数据集,并通过一系列后处理步骤保障图像质量、拓展待分类疾病的覆盖范围,并确保每个疾病标签对应充足的样本量。 MuReD数据集共包含2208张图像,涵盖20种不同标签,图像质量与分辨率存在差异。本研究同时保障了数据的最低质量标准,并为每个标签提供了充足的样本量。据我们所知,MuReD数据集是目前唯一一款通过系列后处理步骤保障图像质量、病理多样性与单标签样本量的公开数据集,可有效提升数据质量并显著降低公开数据集普遍存在的类别不平衡问题。 本研究预期MuReD数据集将助力研发鲁棒性更强、泛化性更佳且可信度更高的视网膜疾病自动检测与分类模型。 文件说明: 1. 训练集数据文件"train_data.csv"包含训练集图像及其对应的20种标签信息。 2. 验证集数据文件"val_data.csv"包含验证集图像及其对应的20种标签信息。 3. "images"文件夹存储组成MuReD数据集的全部图像文件。 该数据集的图像包含两种格式:.tiff与.png。 图像分辨率不统一:由于图像源自不同数据源,其分辨率范围为520×520至3400×2800。



