CAMeL-Lab/BAREC-10M
收藏资源简介:
BAREC-10M Corpus v1.0 是一个扩展的阿拉伯语文本数据集,基于平衡阿拉伯语可读性评估语料库(BAREC)进行扩展,规模从100万词增加到1000万词,并扩展了其覆盖范围,包括平衡的多领域内容。每个文本都标注了领域(如艺术与人文、社会科学或STEM)、读者级别(基础、高级或专业)和文本类别(如教育材料、文学、艺术与音乐、媒体与文化、学术、百科全书或宗教与哲学),并使用先进工具进行了自动形态学、句法和可读性分析。数据集包含文档级(手动标注)和句子级(自动生成)的标注信息,语言为现代标准阿拉伯语。文件结构包括元数据、原始句子文件、形态学与可读性标注文件以及句法标注文件(采用CATiB和UD方案)。数据集适用于文本分类等自然语言处理任务,旨在支持阿拉伯语的可读性研究和多领域分析。
BAREC-10M Corpus v1.0 is an expanded version of the Balanced Arabic Readability Evaluation Corpus (BAREC), scaling from 1 million to 10 million words and broadening its scope to include balanced, multi-domain coverage. Each text is labeled by domain, genre, and readership level, and enriched with automatic morphological, syntactic, and readability analysis using state-of-the-art tools. The corpus includes both document-level (manually labeled) and sentence-level (automatically generated) annotations, with language being Modern Standard Arabic. The dataset structure comprises metadata, raw sentences, morphology and readability annotations, and syntactic annotations in both CATiB and UD schemes. It is designed for text classification and other NLP tasks, supporting Arabic readability research and multi-domain analysis.




