UzMED-ABSA: An Uzbek Medical Domain Dataset for Aspect-Based Sentiment Analysis
收藏资源简介:
UzMED-ABSA is an Uzbek medical-domain dataset for aspect-based sentiment analysis. The dataset contains 7,500 annotated aspect-level rows in TSV format and is designed for research on Uzbek healthcare text analysis, aspect category classification, sentiment classification, polarity scoring, negation detection, and sarcasm or irony detection. Each record includes an Uzbek medical or healthcare-related text sample, aspect term, human-readable aspect category, normalized aspect label, opinion expression, sentiment label, polarity score, intensity label, context type, negation flag, sarcasm flag, script label, and source row number. The dataset covers 35 healthcare aspect categories, including doctor competence, doctor communication, doctor attitude, nurse service, reception process, waiting time, diagnosis accuracy, treatment effectiveness, medication recommendation, laboratory tests, equipment quality, clinic environment, hygiene, emergency care, pain management, surgery quality, inpatient care, billing transparency, insurance benefits, privacy and ethics, patient safety, telemedicine, access, and overall recommendation. The dataset includes 5,725 unique text samples, 5,786 unique sample IDs, and 1,288 unique aspect terms. Sentiment labels are distributed as follows: POS 3,052, NEG 2,940, NEU 1,003, and MIX 505. Polarity scores range from 1 to 5, where 1 indicates strong negative sentiment, 3 indicates neutral or mixed sentiment, and 5 indicates strong positive sentiment. The writing-system distribution is 6,040 Latin-script rows, 801 Cyrillic-script rows, and 659 mixed-script rows. The dataset also includes 824 rows with negation and 202 rows with sarcasm or irony. The TSV file is encoded in UTF-8. During preparation, column names were converted to English snake_case, text fields were Unicode-normalized, whitespace was stripped and collapsed, Uzbek apostrophe variants were normalized to U+02BB, categorical labels were canonicalized, and numeric fields were validated as integers.
UzMED-ABSA是一款面向乌兹别克语医学领域的基于方面的情感分析(Aspect-Based Sentiment Analysis)数据集。该数据集包含7500条带人工标注的方面级条目,格式为TSV,专为乌兹别克语医疗文本分析、方面类别分类、情感分类、极性评分、否定检测以及讽刺/反语检测相关研究设计。 每条记录包含乌兹别克语医疗或健康相关文本样本、方面术语、可读式方面类别、标准化方面标签、意见表达、情感标签、极性得分、强度标签、上下文类型、否定标记、讽刺标记、书写系统标签以及源行编号。该数据集涵盖35个医疗健康方面类别,包括医生能力、医生沟通、医生态度、护士服务、接待流程、等待时长、诊断准确率、治疗有效性、用药推荐、实验室检查、设备质量、诊所环境、卫生状况、急诊护理、疼痛管理、手术质量、住院护理、账单透明度、保险福利、隐私与伦理、患者安全、远程医疗、可及性以及整体推荐度。 该数据集包含5725个独特文本样本、5786个独特样本ID以及1288个独特方面术语。情感标签分布情况如下:正面(POS)3052条、负面(NEG)2940条、中性(NEU)1003条、混合(MIX)505条。极性得分取值范围为1至5,其中1代表强烈负面情感,3代表中性或混合情感,5代表强烈正面情感。书写系统分布为:6040条采用拉丁字母脚本、801条采用西里尔字母脚本、659条采用混合脚本的条目。该数据集还包含824条带否定标记的条目以及202条带讽刺/反语标记的条目。 该TSV文件采用UTF-8编码。在数据制备阶段,列名被转换为英语蛇形命名法,文本字段经过Unicode标准化处理,空白字符被清理并合并,乌兹别克语撇号变体被标准化为U+02BB,分类标签被规范化,数值字段被验证为整数类型。



