遇见数据集

UzThemeLex Dataset: An Uzbek Thematic Lexicon for Domain Terminology and Weakly Supervised NER

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

UzThemeLex is a curated Uzbek-language thematic lexicon dataset designed for domain terminology mining and weakly supervised named entity recognition (NER). The release contains 4,945 unique terminological entries organized into 3 top-level domains (Agronomy, Economics and Business, Law and Governance) and 30 subcategories. Each entry provides the Uzbek term in Latin script, a normalized form for matching, a paraphrased Uzbek definition, domain and subcategory labels, provenance pointers to authoritative sources, and lightweight quality-control signals (heuristic confidence score, review flag, ambiguity flag). Optional fields include aliases and example sentences. The dataset is distributed in multiple formats to support both manual inspection and machine processing. It includes a flat CSV file and a multi-sheet Excel workbook, together with a data dictionary that documents all columns and label sets. For training and pipeline integration, the release also provides JSON/JSONL exports, taxonomy metadata, and ready-to-use pattern files for dictionary-based tagging and weak supervision (e.g., spaCy EntityRuler patterns). A validation script is included to help users verify schema consistency and detect formatting issues (e.g., residual Cyrillic characters and apostrophe normalization). UzThemeLex can be used as (i) a domain dictionary for keyword-based classification and information extraction in Uzbek texts and (ii) a gazetteer for generating weak labels to train or fine-tune NER models. The resource is intended to support Uzbek NLP research and applied text analytics in agriculture, economics, and legal/governance domains.

UzThemeLex是一款经人工甄选的乌兹别克语主题词典数据集,专为领域术语挖掘与弱监督命名实体识别(Named Entity Recognition,NER)任务设计。本次发布包含4945个独特术语条目,分为3个顶级领域(农学、经济学与商学、法律与治理)及30个子类别。每个条目均提供拉丁字母书写的乌兹别克语术语、用于匹配的标准化形式、释义性乌兹别克语定义、领域与子类别标签、指向权威来源的溯源标注,以及轻量化质量控制指标(启发式置信度评分、审核标记、歧义标记)。可选字段还包含别名与示例语句。 本数据集以多种格式发布,以支持人工核查与机器处理。其包含扁平化CSV文件、多工作表Excel工作簿,以及一份用于说明所有字段与标签集的数据字典。为适配模型训练与流程集成,本次发布还提供了JSON/JSONL格式导出文件、分类法元数据,以及可直接使用的模式文件,用于基于词典的标注与弱监督任务(如spaCy EntityRuler模式)。此外还附带了验证脚本,可帮助用户验证数据模式一致性,并检测格式问题(如残留西里尔字符与撇号规范化问题)。 UzThemeLex可作为两类资源使用:其一,作为乌兹别克语文本中基于关键词的分类与信息抽取的领域词典;其二,作为生成弱标签以训练或微调命名实体识别(NER)模型的专名词典。本资源旨在助力乌兹别克语自然语言处理(Natural Language Processing,NLP)研究,以及农业、经济学与法律/治理领域的应用型文本分析。

创建时间:
2026-02-10
二维码
社区交流群
二维码
科研交流群
商业服务