LocalDoc/AzTC-full
收藏资源简介:
AzTC(阿塞拜疆文本语料库)完整版是一个扩展版本的阿塞拜疆语文本语料库,也是阿塞拜疆语中最大的文本语料库之一。该语料库包含约24亿个阿塞拜疆语标记,这些文本是从广泛的来源中编译和清理而来,包括新闻门户、书籍、维基百科和立法文本。文本按文档或段落级别组织,确保每一行是一个连贯的单元,而不是孤立的片段。数据以parquet格式提供,包含text(文档或段落内容)和source(来源标签,指示文本来自哪个集合)字段。
The full version of AzTC (Azerbaijani Text Corpus) is an extended variant of the Azerbaijani text corpus, and ranks among the largest text corpora in the Azerbaijani language. This corpus encompasses approximately 2.4 billion Azerbaijani tokens, with its texts compiled and cleaned from diverse sources including news portals, books, Wikipedia, and legislative documents. The texts are structured at the document or paragraph level, ensuring each line constitutes a coherent unit rather than an isolated fragment. The dataset is distributed in Parquet format, featuring two fields: `text` (content of the document or paragraph) and `source` (a source tag indicating the collection from which the text originates).




