bus_cot_preprocessed
收藏资源简介:
该数据集是一个医学影像数据集,专门用于乳腺影像分析任务。数据集包含6440个训练样本,总大小约648MB。每个样本包含以下核心字段:原始医学图像(image字段)、对应的分割掩码图像(mask字段)、图像文件路径(image_file和mask_file)、图像模态信息(imagemodality)、BI-RADS分类标签(birads)以及推理响应文本(reasoning_response)。数据集适用于医学影像分析、计算机辅助诊断、图像分割、分类任务以及结合视觉与文本的多模态学习研究。BI-RADS标签表明该数据集特别关注乳腺影像报告和数据系统分类,可用于开发自动分类或辅助诊断系统。
This dataset is a medical imaging dataset specifically designed for breast imaging analysis tasks. It contains 6440 training samples with a total size of approximately 648MB. Each sample includes the following core fields: original medical image (image field), corresponding segmentation mask image (mask field), image file paths (image_file and mask_file), image modality information (imagemodality), BI-RADS classification label (birads), and reasoning response text (reasoning_response). The dataset is suitable for medical image analysis, computer-aided diagnosis, image segmentation, classification tasks, and multimodal learning research combining visual and textual information. The BI-RADS label indicates that this dataset focuses particularly on the Breast Imaging Reporting and Data System classification, making it useful for developing automated classification or auxiliary diagnostic systems.
根据您提供的数据集详情页面信息,以下是该数据集的总结:
数据集名称
WOOJYE/bus_cot_preprocessed
数据集地址
https://huggingface.co/datasets/WOOJYE/bus_cot_preprocessed
数据集描述
该数据集是一个经过预处理的乳腺超声(BUS)图像数据集,包含图像、分割掩膜以及相关的推理链(Chain-of-Thought)标注信息,主要用于医学图像分析中的分割与推理任务。
数据集特征
数据集包含以下字段:
| 特征名称 | 数据类型 | 说明 |
|---|---|---|
| id | string | 样本唯一标识符 |
| image | image | 乳腺超声图像 |
| mask | image | 对应的分割掩膜图像 |
| image_file | string | 图像文件名 |
| mask_file | string | 掩膜文件名 |
| imagemodality | string | 影像模态类型 |
| birads | string | BI-RADS分类(乳腺影像报告和数据系统) |
| reasoning_response | string | 推理链响应(Chain-of-Thought推理过程) |
数据集划分
数据集仅包含一个划分:
- 训练集(train):共6,440个样本,数据大小为648,680,162字节(约618.5 MB)
数据集大小
- 下载大小:750,250,981字节(约715.5 MB)
- 数据集总大小:648,680,162字节(约618.5 MB)
配置
- 配置名称:default
- 数据文件路径:
data/train-*




