media-metadata-wikidata-entities
收藏资源简介:
Wikidata Media Entities 是一个从 Wikidata 知识库中通过 SPARQL 查询提取的、专注于媒体行业组织结构的数据集,非通用 Wikidata 数据转储。它精心筛选了与媒体元数据相关的实体,包括公司、工作室和机构,总计约32.6万个条目。数据集覆盖七种核心实体类型:视频游戏开发工作室、电影制作公司、电视网络与流媒体平台、音乐唱片公司、广播电台、图书出版公司以及日本动画工作室。每个实体包含九个结构化字段:Wikidata唯一标识符(QID)、英文名称标签、英文描述、所属国家名称、国家的Wikidata QID、成立年份、解散年份(若活跃则为空)、官方网站URL和实体类型。适用于构建知识图谱、为其他数据集提供信息补充,或作为命名实体识别(NER)任务的训练数据。遵循 CC0 1.0 公共领域贡献许可协议。
Wikidata Media Entities is a dataset extracted from the Wikidata knowledge base through targeted SPARQL queries, focusing on the organizational structure of the media industry. It is not a general Wikidata data dump but a carefully curated collection of entities related to media metadata, including companies, studios, and institutions, totaling approximately 326,000 entries. The dataset covers seven core entity types: video game development studios, film production companies, television networks and streaming platforms, music record labels, radio stations, book publishing companies, and Japanese animation studios. Each entity includes nine structured fields: Wikidata unique identifier (QID), English name label, English description, country of origin, countrys Wikidata QID, founding year, dissolution year (empty if still active), official website URL, and entity type. It is suitable for building knowledge graphs, supplementing information for studio/label fields in other datasets, or serving as training data for named entity recognition (NER) tasks. The dataset follows the CC0 1.0 Public Domain Dedication license.





