遇见数据集

Uzbek Patient Comments for Doctor Specialty Classification

收藏
Mendeley Data2026-08-04 收录
官方服务:

资源简介:

This dataset contains Uzbek patient comments annotated for doctor specialty classification. It is designed for Uzbek medical NLP, text classification, doctor specialty routing, weak supervision analysis, and benchmarking Uzbek language models. The dataset is provided as a UTF-8 TSV file and contains 205,350 master comment records with 15 fields. Each record includes the original patient comment (`text_raw`), a normalized version of the comment (`text_normalized`), the final doctor specialty label (`final_doctor_label`), the corresponding human-readable specialty name (`final_doctor_name`), confidence scores, duplicate-resolution information, manual review indicators, assignment method metadata, and matched signal samples. The main classification target is `final_doctor_label`. The tagset contains 26 labels: ALL, CAR, DEN, DER, EMR, END, ENT, GAS, GYN, HEM, INF, NEU, NONE, NPH, ONC, OPH, PED, PSUR, PSY, PUL, REH, RHE, SUR, TER, TRM, and URO. The `NONE` label is used when the doctor specialty cannot be determined reliably from the comment. The released dataset has no empty values in `text_raw` or `text_normalized`, no duplicate TSV rows, and represents 241,021 original source occurrences through the `occurrence_count` field. The largest classes are `NONE` and `PED`, reflecting the naturally imbalanced distribution of real user-generated patient comments. This dataset is intended for research and development purposes only and should not be used for clinical decision-making.

创建时间:
2026-07-08
二维码
社区交流群
二维码
科研交流群
商业服务