EstNER
收藏资源简介:
EstNER数据集用于爱沙尼亚语的命名实体识别(NER),包含两个部分:'新爱沙尼亚NER数据集'和'重新标注的爱沙尼亚NER数据集'。每个部分进一步分为训练、开发和测试集。数据集包含多达三个层次的嵌套实体的分层标注。标注的实体包括人名、地缘政治实体、地理位置、组织、产品、事件、日期、时间、头衔、货币表达和百分比。README文件还提供了每个数据集部分的统计数据,包括文档数量、句子数量、标记数量和每个层次的实体数量。此外,README文件包含用于引用的BibTeX条目。
The EstNER dataset is designed for Estonian named entity recognition (NER), and it consists of two components: the 'New Estonian NER Dataset' and the 'Relabeled Estonian NER Dataset'. Each component is further split into training, development, and test subsets. The dataset features hierarchical annotations with up to three levels of nested entities. The annotated entity categories include personal names, geopolitical entities, locations, organizations, products, events, dates, times, titles, monetary expressions, and percentages. The accompanying README file provides statistical summaries for each dataset component, including the counts of documents, sentences, tokens, and entities at each annotation level. Additionally, the README includes a BibTeX entry for citation.




