A Large-Scale Annotated Urdu Mental Health Dataset for Anxiety and Depression Classification
收藏资源简介:
This dataset presents the first large-scale annotated Urdu mental health corpus comprising 36,000 tweets collected and constructed for the purpose of automated mental health classification. The dataset is categorised into three classes: anxiety (12,833 tweets), depression (11,617 tweets), and neutral (12,543 tweets). The corpus was constructed through a multi-stage pipeline: (1) manual annotation of 1,000 seed tweets collected via the Twitter API by domain experts; (2) neural machine translation of an English mental health tweet corpus into Urdu using the Helsinki-NLP opus-mt-en-ur model; (3) semi-supervised label propagation using a confidence-threshold classifier trained on the seed set; and (4) a two-stage validation protocol combining human Likert-scale scoring and automated cross-lingual semantic consistency verification. This dataset is intended to support research in Urdu natural language processing, computational mental health detection, and low-resource language modelling. It is released alongside the paper "A Large-Scale Annotated Urdu Corpus and Deep Learning Benchmark for Mental Health Classification" submitted to Scientific Reports (Submission ID: 245269c9-1ee2-467d-bd18-77e14f7b0113).



