GPT-NL Public Corpus
收藏资源简介:
GPT-NL公共语料库是由荷兰应用科学研究组织联合多家机构构建的荷兰语优先大模型预训练数据集,包含21个子集共360亿荷兰语Token及部分英语、代码等多语言数据。该数据集整合了Common Corpus等现有语料库的精选内容,并通过合作机构采集或合成增强技术新增荷兰语数据,所有数据均遵循CC-BY许可。其核心目标是为商业及非商业用途提供合法、低偏见且高质量的语料,支持荷兰语及多语言模型的开发,解决低资源语言训练数据稀缺与版权合规问题。
The GPT-NL Public Corpus is a Dutch-first large language model (LLM) pre-training dataset developed by the Netherlands Organization for Applied Scientific Research in collaboration with multiple institutions. It comprises 21 subsets, totaling 36 billion Dutch Tokens, along with multilingual data including partial English and code datasets. This corpus integrates curated content from existing corpora such as Common Corpus, and adds augmented Dutch language data via collection or synthetic data enhancement technologies from partner institutions. All data is licensed under CC-BY. Its core objective is to provide legal, low-bias, high-quality corpora for both commercial and non-commercial use, supporting the development of Dutch and multilingual models, and addressing the issues of scarce training data and copyright compliance for low-resource languages.
GPT-NL/Collection-metadata 数据集概述
数据集简介
该数据集页面展示了构成GPT-NL语料库的元数据集合。GPT-NL是一个用于大型语言模型预训练的荷兰语语料库。
语料库构成
GPT-NL语料库由公共语料库和私有语料库两部分组成,其数据由多个合作方贡献。
GPT-NL 公共语料库贡献方
- kb
- vng
- officiele-bekendmakingen
- woogle
- Tweede-Kamer
- Rijksoverheid
- Rechtspraak
- Nationaal Archief
- Utrechts Archief
- Noord-Hollands Archief
- Zeeuws Archief
- DANS
- Naturalis
- Wikiwijs
- EP
- CommonCorpus
GPT-NL 私有语料库贡献方
- NDP Nieuwsmedia
- BNR
- NTvG
- ANP
- DNB
- ICTRecht
- ivdnt
- Movisie
- Centerdata
- Waarbenjij
- Iselinge
- Saxion
- Driestar
注意:所有在私有语料库中拥有数据集的合作方均已签署《内容贡献者协议》(可在 https://gpt-nl.nl/samenwerken/content-board 找到)。该协议规定了GPT-NL团队与数据贡献者之间的统一协议和责任。
元数据详情
所有语料库集合(包括公共语料库如 American-stories 和私有语料库如 Instituut voor de Nederlandse Taal)的元数据均可在 GPT-NL-Corpus-metadata.json 文件中查看。对于每个集合,该文件回答了以下几个问题:
- 该数据集是什么?
- 它来自哪里?
- 使用权利是什么?
- 整体数据质量如何?
- 它涵盖什么时间段?
- 它在“数字上”看起来是什么样子?
注意:已填写的元数据是与数据贡献者合作完成的。部分数据集的元数据收集仍在进行中,并可能在收到更多信息后进行调整。
元数据文件示例
文件中为每个集合提供了一个详细的条目。以“荷兰语研究所”为例,条目包含以下字段:
- description: 数据集的描述。
- origin: 数据来源的详细说明。
- modality: 数据模态(例如“text”)。
- license: 许可类型(例如
["GPT-NL Proprietary"])。 - relevance: 相关性评级(如“high”)及其理由。
- quality: 质量评级(如“medium”)及其理由。
- temporal_coverage: 时间覆盖范围分布(例如,1950年前占5%,1950-2000年占75%等)。
- acquisition: 获取方法、收集时间和收集者信息。
- processing: 对数据所做的修改。
- notes: 其他说明,例如是否包含个人身份信息(PII)。




