ceres-open-data-index
收藏资源简介:
Ceres Open Data Index 是一个经过整理和去重的开放数据集索引,包含来自9个国家和国际来源的25个CKAN门户的349,836个开放数据集的元数据。该数据集是目前最大的单一可下载资源形式的聚合开放数据元数据索引。数据集包含政府机构开放数据门户的规范化元数据,通过CKAN API由Ceres(一个开放数据的语义搜索引擎)采集。元数据已从CKAN的嵌套JSON扁平化为表格形式,并进行了噪声过滤和跨门户重复标记。数据集结构包括原始ID、来源门户、门户名称、URL、标题、描述、标签、组织、许可证等字段。数据集以Parquet格式提供,包含完整数据集和按门户划分的子集。适用于文本分类、特征提取等任务,尤其适合开放数据搜索、元数据分析等应用场景。数据集存在地理和语言分布上的偏差,且仅包含CKAN门户的数据。
The Ceres Open Data Index is a curated and deduplicated open dataset index that holds metadata for 349,836 open datasets across 25 CKAN portals from 9 national and international sources. This dataset is the largest aggregated open data metadata index available as a single downloadable resource to date. The dataset contains normalized metadata harvested by Ceres—a semantic search engine dedicated to open data—from government agency open data portals via CKAN APIs. The metadata has been flattened from CKAN's nested JSON format into tabular form, with noise filtering and cross-portal deduplication applied. The dataset structure includes fields such as original ID, source portal, portal name, URL, title, description, tags, organization, license, and others. The dataset is provided in Parquet format, including both the full dataset and portal-partitioned subsets. It is applicable for tasks including text classification and feature extraction, and is particularly suitable for use cases such as open data search and metadata analysis. The dataset exhibits geographic and linguistic distribution biases, and only includes data from CKAN portals.



