遇见数据集

TibNER:Tibetan Named Entity Recognition Dataset

收藏
科学数据银行2024-12-19 更新2026-04-23 收录
官方服务:

资源简介:

Structured linguistic resources are an important foundation for natural language processing. Currently, due to the lack of open-source datasets, the research on Tibetan named entity recognition progresses slowly and the results accumulate less. Based on this, this paper semi-automatically constructs a Tibetan named entity recognition dataset (TibNER) using an entity dictionary. In order to ensure the quality of the dataset, the automatic annotation results are manually proofread.TibNER contains 20,096 sentences, with an average sentence length of 44.2069 syllables, and the annotated entities include names of people, places, and organizations, with a total number of 43,678 in the three types of entities.In order to validate the validity of the dataset, this paper conducts a comparative test on three types of mainstream sequence annotation models, with an F1 value of up to 80.60%. After the study, this data provides data construction experience for low-resource languages, and provides certain data basis for studies such as Tibetan named entity recognition.

创建时间:
2024-02-20
二维码
社区交流群
二维码
科研交流群
商业服务