遇见数据集

ANERD: A Large-Scale Arabic Corpus for Named Entity Recognition and Disambiguation

收藏
Zenodo2026-08-12 更新2026-08-13 收录
官方服务:

资源简介:

ANERD v1.0.0 is a large-scale Arabic corpus designed to support both Named Entity Recognition (NER) and Named Entity Disambiguation/Entity Linking (NED/EL). The corpus contains 502,756 sentences and 11,551,982 tokens. Named entities are annotated using the BIO scheme with four entity types: PERS, ORG, LOC, and MISC. Entity mentions are linked, when applicable, to Arabic Wikipedia through normalized entity titles and URLs; mentions without an appropriate knowledge-base target are marked as NIL. The corpus is organized into Easy, Medium, and Hard sentence-level difficulty categories based on annotation agreement. Each difficulty category is further divided into training, validation, and test subsets using a 60%/20%/20% split. The dataset is distributed as UTF-8 plain-text files in a sentence-delimited, token-per-line format.

提供机构:
Zenodo
创建时间:
2026-08-12
二维码
社区交流群
二维码
科研交流群
商业服务