GPT-NL/GPT-NL_Public_Corpus
收藏资源简介:
GPT-NL公共语料库是最大的宽松许可荷兰语资源,用于大型语言模型预训练。它包含29个精选集合,总计超过5240亿标记,包括360亿荷兰语、2070亿英语、2320亿代码和480亿德语/丹麦语标记。所有数据均基于宽松许可,并以CC-BY许可证重新发布。语料库旨在支持安全、透明和值得信赖的LLM应用,符合欧洲和荷兰的公共价值观。数据来源广泛,包括政府、司法、档案/公共领域、教育/科学和欧盟文件等多个领域,还包括合成和过滤的网络爬取数据。
The GPT-NL Public Corpus is the largest openly licensed Dutch-language resource for large language model (LLM) pre-training. It comprises 29 curated collections, totaling over 524 billion tokens, including 36 billion Dutch, 207 billion English, 232 billion code, and 48 billion German/Danish tokens. All data is sourced under open licenses and republished under the CC-BY license. The corpus is designed to support safe, transparent, and trustworthy LLM applications, aligning with European and Dutch public values. It draws from a wide range of sources, covering multiple domains such as government, judicial, archival/public domain, educational/scientific, and EU documents, as well as synthetic and filtered web-crawled data.




