遇见数据集

Simulated pollen microscope slides for segmentation/bounding box regression and classification

收藏
Zenodo2020-10-09 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Pollen dataset</strong> Sebastian Seurig, Institute for Medical Informatics and Biometry, Carl Gustav Carus Faculty of Medicine, TU Dresden<br> (https://orcid.org/0000-0001-6511-0102)<br> Peter Steinbach, Helmholtz AI, Artificial Intelligence Cooperation Unit, Helmholtz-Zentrum Dresden-Rossendorf<br> (https://orcid.org/0000-0002-4974-230X)<br> Nico Scherf, Institute for Medical Informatics and Biometry, Carl Gustav Carus Faculty of Medicine, TU Dresden and Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig<br> (https://orcid.org/0000-0003-4003-9121)<br> Ingo Röder, Institute for Medical Informatics and Biometry, Carl Gustav Carus Faculty of Medicine, TU Dresden This artificial dataset was created for testing machine learning models on segmentation or bounding box regression and classification tasks.<br> It contains 80.000 (1280x1280 RGB) images of airborne pollen of different sizes within their natural range: 2x10.000 images each Chenopodium bonus-henricus and Corylus colurna making up the medium-sized pollen 10.000 images Urtica dioica making up the smaller pollen 10.000 images Secale cereale making up the larger pollen 4x10.000 images each containing artifacts such as dust or bursted pollen and a mix of pollen in the following ratios:<br> - 'equal_split' 25% each class<br> - 'smaller_pollen_split' 70% Urtica dioica, *<br> - 'middle_pollen_split' 40% Corylus colurna, 10% Chenopodium bonus-henricus, *<br> - 'bigger_pollen_split' 70% Secale cereale, *<br> - * the rest of the pollen split equally between the remaining classes<br> Labels consist of : monochrome masks of each pollen slide (ignoring artifacts) for segmentation x and y-coordinates of the bounding boxes containing all pixels of each of the pollen for regression class names for each labeled pollen. For further questions and suggestions, please do not hesitate to contact the authors. Extensive code and configuration settings for the generator used can be found at https://github.com/seurig/slide-generator.<br> All raw images used to generate this dataset were taken from https://pollen.tstebler.ch/.

**花粉数据集(Pollen dataset)** 作者:Sebastian Seurig,德累斯顿工业大学卡尔·古斯塔夫·卡鲁斯医学院医学信息学与生物统计学研究所(ORCID:https://orcid.org/0000-0001-6511-0102) Peter Steinbach,亥姆霍兹人工智能合作单元(Helmholtz AI)、亥姆霍兹德累斯顿罗森多夫研究中心(ORCID:https://orcid.org/0000-0002-4974-230X) Nico Scherf,德累斯顿工业大学卡尔·古斯塔夫·卡鲁斯医学院医学信息学与生物统计学研究所及莱比锡马克斯·普朗克人类认知与脑科学研究所(ORCID:https://orcid.org/0000-0003-4003-9121) Ingo Röder,德累斯顿工业大学卡尔·古斯塔夫·卡鲁斯医学院医学信息学与生物统计学研究所 本人工数据集专为测试机器学习模型的图像分割、边界框回归与分类任务而构建。数据集共包含80000张分辨率为1280×1280的RGB图像,内容为不同尺寸范围内的气传花粉颗粒: - 2组各10000张图像,分别对应*Chenopodium bonus-henricus*与*Corylus colurna*的中等尺寸花粉; - 10000张图像对应*Urtica dioica*的小型花粉; - 10000张图像对应*Secale cereale*的大型花粉; - 4组各10000张图像,包含粉尘、破裂花粉等干扰物,以及按以下比例混合的花粉样本: 1. 均衡划分(equal_split):每类样本占比25%; 2. 小型花粉划分(smaller_pollen_split):70%为*Urtica dioica*,剩余类别按均等比例分配其余份额; 3. 中型花粉划分(middle_pollen_split):40%为*Corylus colurna*、10%为*Chenopodium bonus-henricus*,剩余类别按均等比例分配其余份额; 4. 大型花粉划分(bigger_pollen_split):70%为*Secale cereale*,剩余类别按均等比例分配其余份额。 数据集标签包含三类信息: 1. 每张图像中每颗花粉的单色掩码(忽略干扰物),用于图像分割任务; 2. 包含单颗花粉所有像素的边界框的x、y坐标,用于边界框回归任务; 3. 每颗标记花粉的类别名称,用于分类任务。 如有进一步问题或建议,请随时联系作者。所用数据集生成器的完整代码与配置参数可访问https://github.com/seurig/slide-generator获取。本数据集生成所用的全部原始图像均来自https://pollen.tstebler.ch/。

提供机构:
Zenodo
创建时间:
2020-10-09
二维码
社区交流群
二维码
科研交流群
商业服务