breast-cancer-database
收藏资源简介:
Today, medical image analysis papers require solid experiments to prove the usefulness of proposed methods. However, experiments are often performed on data selected by the researchers, which may come from different institutions, scanners, and populations. Different evaluation measures may be used, making it difficult to compare the methods. In this paper, we introduce a dataset of 7909 breast cancer histopathology images acquired on 82 patients, which is now publicly available from http://web.inf.ufpr.br/vri/breast-cancer-database. The dataset includes both benign and malignant images. The task associated with this dataset is the automated classification of these images in two classes, which would be a valuable computer-aided diagnosis tool for the clinician. In order to assess the difficulty of this task, we show some preliminary results obtained with state-of-the-art image classification systems. The accuracy ranges from 80% to 85%, showing room for improvement is left. By providing this dataset and a standardized evaluation protocol to the scientific community, we hope to gather researchers in both the medical and the machine learning field to advance toward this clinical application.
当前,医学图像分析领域的学术论文需依托严谨的实验验证以证明所提出方法的实用性。然而,现有相关实验通常采用研究者自主选取的数据集,这类数据可能来自不同机构、扫描仪及受试人群,且不同研究使用的评估指标也存在差异,这使得各方法间的横向对比存在较大困难。本文构建了一个包含7909幅乳腺癌组织病理学图像的数据集,这些图像采集自82名患者,目前可通过公开链接http://web.inf.ufpr.br/vri/breast-cancer-database获取。该数据集涵盖良性与恶性两类病理图像,其配套的研究任务为将这些图像自动划分为上述两类,该自动分类系统可作为一款极具应用价值的计算机辅助诊断工具,为临床医师提供辅助支持。为评估该任务的难度,我们采用当前顶尖的图像分类系统开展了初步实验,实验准确率区间为80%至85%,可见该任务仍存在较大的优化空间。我们期望通过向全球学术共同体公开此数据集与标准化评估流程,吸引医学与机器学习领域的研究者共同参与,推动该临床应用的进一步发展。




