遇见数据集

MindGuard Dataset and Code

收藏
Zenodo2026-05-07 更新2026-05-26 收录
官方服务:

资源简介:

This work presents MINDGUARD, a fine-grained benchmark dataset and experimental framework for detecting mental health misinformation in online peer-support communities. The study focuses on Reddit-based mental health discussions, where users often share personal experiences, emotional support, treatment opinions, and coping advice. While such communities can provide valuable support, they may also spread harmful narratives such as anti-medication claims, therapy discouragement, miracle-cure promotion, and symptom minimization. The main aim of this work is to identify and classify these harmful narratives using natural language processing and machine learning techniques. The work develops a manually annotated dataset of 10,000 Reddit posts, categorized into five labels: anti-medication, therapy-discouragement, miracle-cure promotion, symptom-minimization, and other/non-misinformation. This fine-grained taxonomy allows the study to go beyond simple binary misinformation detection and capture different forms of harmful mental health advice. The dataset includes post content along with contextual metadata such as subreddit, author identifier, and timestamp, enabling future analysis of community-level and temporal misinformation patterns. The methodology evaluates both traditional machine learning models and transformer-based language models. Classical models such as Logistic Regression, Naive Bayes, and Linear SVM are trained using TF-IDF features, while transformer models such as BERT, DistilBERT, and RoBERTa are fine-tuned for both binary and fine-grained classification tasks. The architecture diagram on page 5 illustrates the complete workflow, beginning with Reddit data ingestion and preprocessing, followed by feature engineering, model learning through traditional and transformer-based paths, and final inference for binary and multi-class misinformation detection. The experimental results show that machine learning models can effectively detect mental health misinformation. Linear SVM achieved strong performance among classical models, while transformer-based models, especially RoBERTa-base, achieved the best fine-grained classification performance. The results also show that explicit misinformation, such as anti-medication rhetoric and therapy discouragement, is easier to detect because it contains strong lexical signals. In contrast, symptom minimization and miracle-cure narratives are more difficult because they are often written in subtle, emotional, or conversational language. Overall, this work contributes a new dataset, a mental-health-specific misinformation taxonomy, benchmark experiments, and analysis of category-wise detection challenges. It supports future research on safer online peer-support communities by helping moderators, researchers, and platform designers identify harmful mental health advice more accurately and contextually.

提供机构:
Zenodo
创建时间:
2026-05-07
二维码
社区交流群
二维码
科研交流群
商业服务