遇见数据集

Curated BUSI dataset - Curated Breast Ultrasound Images

收藏
Zenodo2026-03-20 更新2026-05-26 收录
官方服务:

资源简介:

The Curated BUSI dataset is a post-processed version presented in [2] of the original Breast Ultrasound Images (BUSI) dataset [1]. The original BUSI dataset consists of 780 breast ultrasound images collected from female patients aged 25–75 years, annotated for three classes (normal, benign, malignant) with corresponding segmentation masks for the lesions. Although widely used, the original BUSI dataset contains significant data quality issues, including duplicate images, inconsistent class labels, and annotation discrepancies, which can bias model training and evaluation. To address these issues, Aumente-Maestro et al. (2025) curated the BUSI dataset by applying an automated duplicate detection algorithm followed by manual review to remove repeated and problematic cases. The resulting Curated BUSI dataset contains images with more reliable labels and segmentation masks, reducing the bias present in the original dataset and enabling more robust development and evaluation of machine learning models for breast ultrasound analysis. References[1] Al-Dhabyani, W., Gomaa, M., Khaled, H., & Fahmy, A. (2020). Dataset of breast ultrasound images. Data in Brief, 28:104863. DOI:10.1016/j.dib.2023.109247.[2] Aumente-Maestro, C., Díez, J., & Remeseiro, B. (2025). A multi-task framework for breast cancer segmentation and classification in ultrasound imaging. Computer methods and programs in biomedicine, 260, 108540. DOI:10.1016/j.cmpb.2024.108540.

提供机构:
Zenodo
创建时间:
2026-03-19
搜集汇总
数据集介绍
Curated BUSI dataset - Curated Breast Ultrasound Images 数据集图片
背景与挑战
背景概述
Curated BUSI数据集是对原始乳腺超声图像数据集的精炼版本,通过去除重复和有问题的样本,提供了更可靠的三类标签(正常、良性、恶性)及对应的分割掩码,旨在消除数据偏差,提升机器学习模型在乳腺超声分析中的鲁棒性。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务