imageomics/TreeOfLife-200M
收藏资源简介:
TreeOfLife-200M是一个包含近2.14亿张图片、代表952,257个物种的计算机视觉模型训练数据集。它结合了来自GBIF、EOL、BIOSCAN-5M和FathomNet四个核心生物多样性数据提供商的图像和元数据,是迄今为止最大的、最多样化的公共机器学习就绪型数据集。该数据集还增加了图像上下文的多样性,包括博物馆标本、相机陷阱和公民科学图像。数据集的严格筛选过程确保每个图像都有尽可能具体的分类标签,为训练BioCLIP 2和未来的生物基础模型提供了一个全面的基础。
TreeOfLife-200M is a computer vision model training dataset with nearly 214 million images representing 952,257 taxa. It combines images and metadata from four core biodiversity data providers: GBIF, EOL, BIOSCAN-5M, and FathomNet. This is the largest and most diverse public ML-ready dataset for computer vision models in biology at release. The dataset also increases image context diversity with museum specimen, camera trap, and citizen science images well-represented. A rigorous curation process ensures each image has the most specific taxonomic label possible, providing a well-rounded foundation for training BioCLIP 2 and future biology foundation models.




