oreilly-animals
收藏资源简介:
The OReilly Animal Menagerie 是一个多模态数据集,收录了来自OReilly Media图书封面上的动物图像及相关元数据。数据集包含1414个样本,每个样本对应一本OReilly图书及其封面上的动物。数据内容涵盖两个核心图像模态:图书封面图像(cover_image)和动物雕刻图像(engraving_image),并附有丰富的结构化元数据。元数据分为三大类:1) 图书信息,包括书名、作者、出版日期、主题分类和图书URL;2) 动物信息,包括动物名称、清洁名称、动物类别、维基数据描述、科学名称、IUCN保护状态等;3) 详细的生物分类学信息,涵盖分类等级、属、种、分类学权威、亚种划分、化石范围等字段,以及来自维基共享资源的图像链接和标题说明。此外,数据集还包含维基百科条目标题、Wikidata QID标识符以及数据质量标记(如需要审查)。数据集整合了OReilly的动物图鉴、维基媒体项目(CC BY-SA 4.0许可)和互联网档案馆的图书元数据,适用于计算机视觉、多模态学习、信息检索、数字人文以及动物分类学相关的研究与教育项目。
The O'Reilly Animal Menagerie is a multimodal dataset that collects animal images and associated metadata from the book covers of O'Reilly Media. It contains 1,414 samples, each corresponding to one O'Reilly book and the animal featured on its cover. The dataset covers two core image modalities: cover_image and engraving_image, and is accompanied by rich structured metadata. The metadata is divided into three categories: 1) Book information, including book title, author, publication date, subject category, and book URL; 2) Animal information, including animal name, cleaned name, animal category, Wikidata description, scientific name, IUCN conservation status, etc.; 3) Detailed taxonomic information, covering fields such as taxonomic rank, genus, species, taxonomic authority, subspecies division, fossil range, as well as image links and caption descriptions from Wikimedia Commons. Additionally, the dataset includes Wikipedia article titles, Wikidata QID identifiers, and data quality markers (e.g., requires review). The dataset integrates resources from O'Reilly's animal illustrated reference, Wikimedia projects (licensed under CC BY-SA 4.0), and book metadata from the Internet Archive, and is applicable to research and educational projects related to computer vision, multimodal learning, information retrieval, digital humanities, and animal taxonomy.
数据集概述:The OReilly Animal Menagerie
基本信息
- 数据集名称: The OReilly Animal Menagerie
- 许可证: CC BY-SA 4.0
- 语言: 英语
- 规模: 1,000 至 10,000 行(实际为 1,414 行)
- 标签: oreilly, books, animals
数据内容
- 来源: 所有数据收集自 OReilly Menagerie,并丰富了元数据。
- 数据格式: 每个数据行以 JSON 格式呈现,包含以下字段:
- 动物信息: 动物名称(
animal_name)、学名(scientific_name)、动物分类(如纲、目、科等) - 书籍信息: 书名(
book_title)、作者(author) - 保护状态: IUCN 状态代码(
iucn_code)、标准(iucn_criteria)、评估年份(iucn_year)、IUCN 页面链接(iucn_url) - 生态数据: 种群趋势(
population_trend)、系统(systems)、栖息地(habitats)、威胁(threats)、保护行动(conservation_actions) - 文档信息: 分布范围(
doc_range)、威胁描述(doc_threats) - 维基媒体信息: 维基百科标题(
wiki_title)、维基数据 ID(qid)、图片链接(commons_photo_url)、分布图链接(commons_range_map_url)
- 动物信息: 动物名称(
许可与用途
- 数据集仅供研究和教育用途。
- 书籍封面图像和动物插图归 OReilly Media, Inc. 所有。
- 元数据来自维基媒体项目(仅元数据部分适用 CC BY-SA 4.0 许可证)和 Internet Archive 的书籍元数据。
- 本数据集与 OReilly Media, Inc. 无关联、未经其认可或赞助。
附加信息
- 数据集详情页地址:https://huggingface.co/datasets/christopher/oreilly-animals
- 关于 OReilly 动物历史的文章可参见:https://www.oreilly.com/content/a-short-history-of-the-oreilly-animals/





