遇见数据集

renilthomas82/renil-wikipedia-test

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含从Wikimedia Enterprise结构化内容API生成的结构化维基百科文章快照。数据集以Parquet格式分发,采用统一模式,优化了分析和机器学习工作负载。当前版本包括英语和法语维基百科的文章命名空间。数据集汇总:预解析的英语和法语维基百科文章,使用Wikimedia Enterprise快照API提取。该数据集包含英语和法语维基百科版本的所有文章,预解析并输出为具有一致模式的结构化数据。数据集以Parquet格式提供,优化了高性能分析查询和高效存储。此版本在所有文件中使用统一的固定模式,使其与DuckDB、pandas、Polars和Apache Spark开箱即用兼容。数据集的新功能包括:解析的参考文献和引用,将维基百科的知识与其真实来源连接;解析的表格,这是维基百科页面信息最密集的部分之一;可信度信号(例如referenceneed和referencerisk),指示信息可能未得到充分来源支持的情况;列表解析的改进,包括嵌套列表、有序列表和定义列表;文章正文图像已解析并在文章正文部分负载中可用。

This dataset contains structured Wikipedia article snapshots generated from the Wikimedia Enterprise Structured Contents API. The dataset is distributed in Parquet format with a unified schema optimized for analytical and machine learning workloads. The release currently includes English and French Wikipedia article namespaces. Dataset Summary: Pre-parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre-parsed and output as structured data with a consistent schema. The dataset is provided in Parquet format, optimized for high-performance analytical queries and efficient storage. This version uses a unified, pinned schema across all files, making it compatible with DuckDB, pandas, Polars, and Apache Spark out of the box. New in this dataset: Parsed references and citations, connecting Wikipedias knowledge with its sources of truth. Parsed tables, one of the most information-heavy sections of Wikipedia pages. Credibility signals, for example referenceneed and referencerisk, signaling where information may not be sufficiently backed by sources. Improvements to list parsing, including nested lists, ordered lists, and definition lists. Article-body images are parsed and available in the article body section payload.

提供机构:
renilthomas82
二维码
社区交流群
二维码
科研交流群
商业服务