遇见数据集

CO.PRE.PAN Full Corpus (Restricted)

收藏
Zenodo2026-02-23 更新2026-05-26 收录
官方服务:

资源简介:

This record contains the complete CO.PRE.PAN (Corpus de Prensa Panhispánico) press corpus, organized into country-specific ZIP archives with linguistically annotated JSON files. Due to copyright restrictions, all texts and annotations are distributed under restricted access and cannot be shared openly. Users may request access directly through Zenodo. Contents of this record Each {COUNTRYCODE}.zip archive contains: Press texts in plain text format (txt-files/) Annotated JSON files (json-annotated/) All ZIP archives were generated using the internal script "zenodo_corpus_zip.py", which automatically tracks timestamps and file changes to ensure reproducible versioning. Corpus description CO.PRE.PAN is a cross-national corpus of written press Spanish from 18 Spanish-speaking countries, comprising over 14 million words. It is structurally aligned with the spoken broadcast corpus CO.RA.PAN (Corpus Radiofónico Panhispánico) and serves as a scripted register baseline for comparative analyses of national standard varieties of Spanish. All texts are drawn from comparable press genres and produced under broadly equivalent publication conditions across countries, ensuring cross-national comparability. Versioning Each version of this record represents a coherent snapshot of the full corpus at a specific point in time. Updates may include newly added texts, corrected or extended annotations, and improvements to preprocessing and linguistic annotation. Annotation details Each JSON file contains: tokenization, sentence segmentation POS tags, lemmas, and morphological features dependency relations automatic categorization of verbal tense and related features All annotations are generated using spaCy (model: es_dep_news_trf), followed by project-specific quality control steps, using the same annotation pipeline applied to CO.RA.PAN. Legal and access information The restricted status of this record is due to copyright limitations. Only short text extracts may be displayed publicly under scientific quotation rules and text-and-data-mining provisions of EU Directive 2019/790 and the German UrhG (§51, §60d, §44b). Redistribution or reuse of the full texts and annotations is not permitted. Access requests can be submitted directly through Zenodo. For scientific inquiries or technical questions, please contact the CO.PRE.PAN project team.

本数据集包含完整的**CO.PRE.PAN(Corpus de Prensa Panhispánico,泛伊比利亚报语文本语料库)**,其按国家分类打包为ZIP压缩档案,内含经过语言学标注的JSON文件。鉴于版权限制,本语料库的所有文本与标注文件均采用受限访问机制,不得公开分享。用户可通过Zenodo平台直接提交访问申请。 ## 数据集内容 每个以国家代码({COUNTRYCODE})命名的ZIP压缩档案包含以下内容: - 纯文本格式的报语文本(存放于txt-files/目录下) - 经过语言学标注的JSON文件(存放于json-annotated/目录下) 所有ZIP压缩档案均通过内部脚本`zenodo_corpus_zip.py`生成,该脚本可自动追踪时间戳与文件变更,保障版本的可复现性。 ## 语料库概况 CO.PRE.PAN是覆盖18个西班牙语国家的跨国民间西班牙语报语文本语料库,总词量逾1400万。其结构与广播口语语料库**CO.RA.PAN(Corpus Radiofónico Panhispánico,泛伊比利亚广播语料库)**保持一致,可作为西班牙语国家标准变体对比研究的书面语基准语料库。所有文本均选自可比的报刊体裁,且在各国基本一致的出版条件下生成,确保了跨国家的可比性。 ## 版本管理 本数据集的每个版本均为语料库在特定时间点的完整快照。更新内容可能包括新增文本、修正或扩展的标注文件,以及预处理与语言学标注流程的优化。 ## 标注详情 每个JSON文件包含以下标注信息: - 分词与分句标注 - 词性标注、词形还原与形态特征标注 - 依存关系标注 - 动词时态及相关特征的自动分类标注 所有标注均通过spaCy工具(模型为es_dep_news_trf)生成,随后经过项目专属的质量管控流程,所用标注流程与CO.RA.PAN保持一致。 ## 法律与访问说明 本数据集采用受限访问模式源于版权限制。仅可按照科学引用规则,以及欧盟指令2019/790与德国著作权法(UrhG §51、§60d、§44b)中的文本与数据挖掘相关条款,公开展示少量文本摘录。不得对完整文本与标注文件进行重新分发或二次使用。 访问申请可直接通过Zenodo平台提交。若有学术咨询或技术问题,请联系CO.PRE.PAN项目团队。

提供机构:
Zenodo
创建时间:
2026-02-23
二维码
社区交流群
二维码
科研交流群
商业服务