遇见数据集

AraSeg-2026-Shared-Task-PA

收藏
魔搭社区2026-05-29 更新2026-07-19 收录
官方服务:

资源简介:

# Arabic Sentence Segmentation Shared Task 2026 For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit: https://www.araseg.aramlab.ai/ ## Dataset Summary AraSeg is the first comprehensive benchmark for Arabic sentence segmentation. The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy. AraSeg contains manually annotated documents collected from diverse sources and genres, enabling robust evaluation across different writing styles and domains. AraSeg-PA is the Paragraph-Aware (PA) variant of the corpus where documents include paragraph boundaries and punctuation marks. --- ## Dataset Structure ### Data Instances ``` {'doc_id': 'doc_00b450a96684', 'tokens': ['الفصل','الأول','حين','ركبت','السيارة','لم','أكن','أتصور','أنني','أبدأ', ...], 'labels': [0, 1, 0, 0, 0, 0, 0, 0, 0, 0, ...], 'text': 'أبدأ الفصل الأول حين ركبت السيارة لم أكن أتصور أنني...', 'label_str': '0100000000...'} ``` ### Data Fields - **doc_id**: Unique document identifier. - **tokens**: White-space-tokenized document represented as a list of tokens. - **labels**: Token-level sentence boundary labels. `1` indicates that a sentence boundary follows the current token, while `0` indicates no boundary. - **text**: Document. - **label_str**: Sentence boundary labels as a binary string. `1` indicates that a sentence boundary follows the current token, while `0` indicates no boundary. ### Data Splits - **train**: 174 documents (10,657 sentences and 128K words). - **dev**: 222 documents (12,985 sentences and 164K words). - **test**: 262 documents (12,509 sentences and 159K words). --- ## Task Definition Sentence segmentation is formulated as a binary token classification task. Given a sequence of tokens: ```[token_1, token_2, ..., token_n]```, the model predicts for each token whether a sentence boundary follows it. For example: | Token | Label | |---|---| | ذهب | 0 | | الطالب | 0 | | إلى | 0 | | المدرسة | 1 | The label `1` indicates that the sentence ends after the token. --- ## Evaluation We evaluate systems using boundary-level metrics: - **Boundary Precision (P):** Percentage of predicted sentence boundaries that are correct. - **Boundary Recall (R):** Percentage of gold sentence boundaries correctly identified. - **Boundary F1 (F1):** Harmonic mean of precision and recall. Metrics are computed at the document level and averaged across the corpus. We provide evaluation scripts on this [repo](https://github.com/mbzuai-nlp/araseg-shared-task-2026). ---

提供机构:
maas
创建时间:
2026-05-20
二维码
社区交流群
二维码
科研交流群
商业服务