遇见数据集

UzNER-5Style: A 75,000-Sentence Multi-Style Uzbek Named Entity Recognition Dataset

收藏
Mendeley Data2026-09-09 收录
官方服务:

资源简介:

UzNER-5Style is a large-scale Uzbek named entity recognition dataset created across five functional styles: official, journalistic, scientific, literary, and conversational. The dataset contains 75,000 sentences, with 15,000 sentences for each style. The corpus is annotated using the BIOES tagging scheme and includes 30 named entity types. These include person names, organizations, geopolitical entities, dates, monetary values, quantities, durations, percentages, locations, events, laws, documents, positions, phone numbers, e-mail addresses, URLs, times, titles, works of art, academic degrees, and several other entity categories. The final dataset contains approximately 750,000 token-level records. Each record includes a token identifier, source information, functional style, sentence identifier, token, BIOES tag, domain information, and validation status. The dataset was synthetically generated and then subjected to several quality-control stages. These included duplicate checking, BIOES sequence validation, entity-boundary correction, semantic error correction, and text-level cleaning. A subset of 5,000 sentences, consisting of 1,000 sentences from each functional style, was independently reviewed by two annotators. Disagreements were subsequently resolved to obtain final adjudicated labels for the validation subset. This dataset is intended to support research on Uzbek named entity recognition, low-resource natural language processing, cross-style generalization, domain-aware NER, and robustness across different functional styles.

创建时间:
2026-08-26
二维码
社区交流群
二维码
科研交流群
商业服务