遇见数据集

Catalan Government Crawling

收藏
Zenodo2024-02-28 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39.117.909 tokens, 1.565.433 sentences and 71.043 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus. We license the actual packaging of this data under a CC0 1.0 Universal License.

提供机构:
Zenodo
创建时间:
2021-03-25
二维码
社区交流群
二维码
科研交流群
商业服务