遇见数据集

ArmTDP-NER: Eastern Armenian Named Entity Corpus

收藏
Zenodo2025-12-20 更新2026-05-26 收录
官方服务:

资源简介:

This gold-standard corpus comprises approximately 150,000 tokens of Modern Eastern Armenian, featuring high-accuracy manual annotations of named entities. Derived from 1,949 news articles published between 2018 and 2020, the dataset covers a diverse range of domains. The system utilizes a modified OntoNotes 5.0 markup scheme, adopting an 18-type superset of the ACE name guidelines. With 17,632 annotated entities, this corpus serves as a robust benchmark for both Entity Detection and Recognition (EDR) and Named Entity Recognition (NER) tasks in Armenian NLP. Entity type Description Count PERSON People including fictional 3,502 NORP Nationalities or religious or political groups 823 FACILITY Buildings, airports, highways, bridges, etc. 601 ORGANIZATION Companies, agencies, institutions, etc. 3,388 GPE Countries, cities, states 3,290 LOCATION Non-GPE locations, mountain ranges, bodies of water 166 PRODUCT Vehicles, weapons, foods, etc. (Not services) 107 EVENT Named hurricanes, battles, wars, sports events, etc. 211 WORK OF ART Titles of books, songs, etc. 156 LAW Named documents made into laws 466 LANGUAGE Any named language 47 DATE Absolute or relative dates or periods 1,995 TIME Times smaller than a day 405 PERCENT Percentage (including "%") 112 MONEY Monetary values, including unit 267 QUANTITY Measurements, as of weight or distance 343 ORDINAL "first", "second" 608 CARDINAL Numerals that do not fall under another type 1,145 Furthermore, a pre-trained NER model based on this dataset—developed by Shake Hakobyan—is available as part of the Stanza NER models, achieving an F1 score of 87.96.

提供机构:
Zenodo
创建时间:
2025-12-20
二维码
社区交流群
二维码
科研交流群
商业服务