project-oceania/planktonzilla-17M
收藏资源简介:
Planktonzilla-17M 是一个大规模、全面的数据集,整合了来自所有公开可用、已标记浮游生物数据集的1700万张浮游生物图像。该统一集合使研究人员能够训练鲁棒的深度学习模型,用于在不同成像系统和海洋学环境中进行浮游生物识别和分类。数据集整合了来自FlowCAM、ISIIS、UVP5/UVP6、ZooScan、ZooCAM、PlanktoScope、IFCB以及无透镜显微镜系统的图像,为浮游生物研究提供了分类学上一致且可用于研究的资源。它是Inria Challenge OcéanIA倡议的一部分。每张图像都包含标准化的分类学层次结构(界、门、纲、目、科、属、种),并映射到世界海洋物种名录(WoRMS),以及丰富的元数据,包括浮游生物分类(活体浮游生物与非浮游生物)和生存状态(活体生物与碎屑/惰性物)的布尔指标,以及捕获采集位置、深度、采样条件和环境背景的地理环境元数据。保留了与源数据集和原始标签的完全可追溯性,以确保完全透明。迄今为止,Planktonzilla-17M代表了有史以来组装的最大、分类学最完整的浮游生物数据集,涵盖了所有主要分类群中的601个不同浮游生物类别,以及具有完整分类谱系的201个已知物种。数据集捕获了从微浮游生物(20微米)到大型浮游动物(>300微米)的生物体,采集时间从2008年到2025年,地点涵盖全球海洋、沿海区域和淡水环境。
Planktonzilla-17M is a large-scale, comprehensive dataset combining 17 million plankton images from all publicly available, labeled plankton datasets. This unified collection enables researchers to train robust deep learning models for plankton identification and classification across diverse imaging systems and oceanographic environments. The dataset integrates imagery from FlowCAM, ISIIS, UVP5/UVP6, ZooScan, ZooCAM, PlanktoScope, IFCB, and lensless microscopy systems, providing a taxonomically consistent and research-ready resource for plankton research. It is part of the Inria Challenge OcéanIA initiative. Each image includes a standardized taxonomic hierarchy (Kingdom, Phylum, Class, Order, Family, Genus, Species) mapped to the World Register of Marine Species (WoRMS), along with rich metadata including boolean indicators for plankton classification (live plankton vs. non-plankton) and living status (living organisms vs. detritus/inert), as well as geo-environmental metadata capturing collection location, depth, sampling conditions, and environmental context. Complete traceability to source datasets and original labels is preserved for full transparency. To date, Planktonzilla-17M represents the largest and most taxonomically complete plankton dataset ever assembled, encompassing 601 distinct plankton classes across all major taxonomic groups and covering 201 known species with complete taxonomic lineage. The dataset captures organisms ranging from microplankton (20 μm) to large zooplankton (>300 μm), collected from 2008-2025 across global oceans, coastal regions, and freshwater environments.




