遇见数据集

LEXICON-BASED SEMANTIC TAGGING INFORMATION SYSTEM FOR THE UZBEK LANGUAGE CORPUS

收藏
Zenodo2026-05-11 更新2026-05-26 收录
官方服务:

资源简介:

This study presents a lexicon-based semantic tagging information system developed for the Uzbek language corpus. The system employs the six-volume Explanatory Dictionary of the Uzbek Language (OʻzTIL) as its primary lexical resource, which contains over 85,000 entries with full semantic definitions, making it the most authoritative normative lexicographic source for Uzbek. An ontological model organized in three hierarchical levels – top, mid, and low – was designed to categorize lexical units extracted from the dictionary. Five core semantic categories were formed: animal names (approximately 100–200 units), bird names (approximately 100–150 units), personal nouns (approximately 500+ units), place names (approximately 300+ units), and occupation names (approximately 200+ units), totaling approximately 1,200–1,400 lexical units. A rule-based automatic tagging algorithm was developed to annotate corpus tokens against this structured lexical database, assigning standardized semantic tags. The system addresses key challenges inherent to Uzbek, including agglutinative morphology and lexical ambiguity. Compared to international systems such as WordNet and USAS, the proposed dictionary-based approach demonstrates superior normative grounding and cultural adequacy for Uzbek. The system is intended to serve as a foundational open resource for downstream natural language processing tasks, including machine translation, information retrieval, and intelligent educational applications.

提供机构:
Zenodo
创建时间:
2026-05-11
二维码
社区交流群
二维码
科研交流群
商业服务