media-metadata-openlibrary-books
收藏资源简介:
TigreGotico/media-metadata-openlibrary-books 是一个由 metadatarr 爬虫工具从 OpenLibrary 抓取并构建的丰富实体数据集。数据集包含约 299 万条图书元数据记录,旨在提供结构化的图书信息用于媒体元数据相关的分析与应用。数据内容涵盖广泛的图书属性,具体包括:图书唯一标识符 (olid)、标题、副标题、作者列表、作者标识符 (author_key)、首次出版年份、主题分类、ISBN-10 与 ISBN-13 编码、出版社、语言、页数中位数、电子书访问权限状态 (ebook_access)、是否拥有全文 (has_fulltext)、版本数量 (edition_count) 以及封面图片标识符 (cover_i)。该数据集适用于多种下游任务,例如图书信息检索、知识图谱实体填充、推荐系统特征工程、出版趋势分析以及数字图书馆资源编目。
TigreGotico/media-metadata-openlibrary-books is a rich entity dataset built by the metadatarr crawler tool, scraped from OpenLibrary. It contains approximately 2.99 million book metadata records, aiming to provide structured book information for media metadata-related analysis and applications. The data covers a wide range of book attributes, including: book unique identifier (olid), title, subtitle, author list, author identifier (author_key), first publication year, subject classifications, ISBN-10 and ISBN-13 codes, publisher, language, median page count, ebook access status (ebook_access), full-text availability (has_fulltext), edition count (edition_count), and cover image identifier (cover_i). This dataset is suitable for various downstream tasks, such as book information retrieval, knowledge graph entity population, feature engineering for recommendation systems, publishing trend analysis, and digital library resource cataloging.




