LocalDoc/azerbaijani-pretrain-corpus
收藏资源简介:
一个清洗过的阿塞拜疆语文本语料库,专为语言模型预训练而构建,合并了两个精选来源并移除了完全重复的文本。内容包括约6,931,898个文档和约53.6亿个令牌(使用`o200k_base`分词器测量;阿塞拜疆语特定分词器会产生更少的令牌,因为`o200k_base`对阿塞拜疆语的粘着性分词效率较低)。每个文档平均约773个令牌。字段包括`text`(文档或段落)和`source`(来源标签:`aztc`或`oscar`)。构建过程涉及合并AzTC-full(新闻、书籍、维基百科、立法等精选语料库)和community_oscar_azerbaijani_scored(基于OSCAR的网页语料库,仅保留质量分数≥2.5的文档),然后通过规范化文本的哈希值移除完全重复项(重复率约28%)。注意:仅移除完全重复项,近重复项未处理;OSCAR部分通过模型生成的质量分数过滤,但未经验证;原始网页文本可能包含个人数据,未进行PII移除;AzTC和OSCAR部分在清洁度上有所不同,可使用`source`字段进行加权或分离。来源数据集包括LocalDoc/AzTC-full和LocalDoc/community_oscar_azerbaijani_scored,需遵守相关许可证。
A cleaned Azerbaijani text corpus assembled for language-model pretraining, merging two curated sources and removing exact duplicates. It contains approximately 6,931,898 documents and ~5.36B tokens (measured with the `o200k_base` tokenizer; an Azerbaijani-specific tokenizer will yield fewer tokens, as `o200k_base` segments agglutinative Azerbaijani inefficiently), with an average of ~773 tokens per document. Fields include `text` (the document or passage) and `source` (origin label: `aztc` or `oscar`). The corpus was built by merging AzTC-full (curated Azerbaijani corpus including news, books, Wikipedia, legislation) and community_oscar_azerbaijani_scored (OSCAR-derived web corpus with quality scores, only documents with quality score >= 2.5 were kept), followed by exact deduplication using hashes of normalized text (NFKC, whitespace-collapsed, casefolded), removing about 28% duplicates. Note that only exact duplicates were removed, near-duplicates remain; the OSCAR portion was filtered by a model-generated quality score (threshold 2.5), though not verified against human annotation; raw web text may contain personal data with no PII removal applied; AzTC and OSCAR portions differ in cleanliness, and the `source` field can be used for weighting or separation. Source datasets are LocalDoc/AzTC-full and LocalDoc/community_oscar_azerbaijani_scored, with licenses to be reviewed and aligned.




