MBZUAI/AraSeg-2026-Shared-Task-NoPnx-NP
收藏资源简介:
AraSeg是第一个用于阿拉伯语句子分割的综合基准数据集。该语料库旨在支持现代标准阿拉伯语(MSA)中的句子分割研究,特别是在标点符号不一致、缺失或嘈杂的环境中。AraSeg包含从多种来源和体裁收集的手动标注文档,能够实现跨不同写作风格和领域的鲁棒评估。AraSeg-NoPnx-NP是该语料库的无标点无段落(NoPnx-NP)变体,其中文档不包含段落边界或标点符号。数据集结构包括文档ID、标记化令牌、句子边界标签、原始文本和标签字符串等字段,并分为训练集(174个文档)、开发集(222个文档)和测试集(262个文档)。任务定义为二元标记分类,即预测每个令牌后是否跟随句子边界。评估采用边界级别的精确率、召回率和F1分数指标。
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation. The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy. AraSeg contains manually annotated documents collected from diverse sources and genres, enabling robust evaluation across different writing styles and domains. AraSeg-NoPnx-NP is the No-Punctuation No-Paragraph (NoPnx-NP) variant of the corpus where documents do not include paragraph boundaries or punctuation marks. The dataset structure includes fields such as doc_id, tokens, labels, text, and label_str, with splits into train (174 documents), dev (222 documents), and test (262 documents). The task is formulated as a binary token classification to predict whether a sentence boundary follows each token. Evaluation uses boundary-level precision, recall, and F1 metrics.



