Explanation of features.
收藏资源简介:
Context and background. Depression has affected millions of people worldwide and has become one of the most common mental disorders. Early mental disorder detection can reduce costs for public health agencies and prevent other major comorbidities. Additionally, the shortage of specialized personnel is very concerning since depression diagnosis is highly dependent on expert professionals and is time-consuming. Research problems. Recent research has evidenced that machine learning (ML) and natural language processing (NLP) tools and techniques have significantly benefited the diagnosis of depression. However, there are still several challenges in the assessment of depression detection approaches in which other conditions such as post-traumatic stress disorder (PTSD) are present. These challenges include assessing alternatives in terms of data cleaning and pre-processing techniques, feature selection, and appropriate ML classification algorithms. Purpose of the study. This paper tackles such an assessment based on a case study that compares different ML classifiers, specifically in terms of data cleaning and pre-processing, feature selection, parameter setting, and model choices. Methodology. The experimental case study is based on the Distress Analysis Interview Corpus - Wizard-of-Oz (DAIC-WOZ) dataset, which is designed to support the diagnosis of mental disorders such as depression, anxiety, and PTSD. Major findings. Besides the assessment of alternative techniques, we were able to build models with accuracy levels around 84% with Random Forest and XGBoost models, which is significantly higher than the results from the comparable literature which presented the level of accuracy of 72% from the SVM model. Conclusions. More comprehensive assessments of ML classification algorithms and NLP techniques for depression detection can advance the state of the art in terms of improved experimental settings and performance.
研究背景与动因:抑郁症已波及全球数百万人群,成为最为高发的精神障碍之一。早期精神障碍筛查可有效降低公共卫生机构的运营成本,并预防其他严重共病的发生。此外,专业精神卫生人才短缺问题亟待关注——抑郁症诊断高度依赖专业临床人员,且流程耗时耗力。 研究问题:现有研究证实,机器学习(ML)与自然语言处理(NLP)相关工具及技术已为抑郁症诊断带来显著增益,但现有抑郁症检测方法的评估仍存在诸多挑战,尤其当患者同时伴随创伤后应激障碍(PTSD)等其他精神疾病时。此类挑战涵盖数据清洗与预处理技术、特征选择方法,以及适配的机器学习分类算法等维度的方案选型评估。 研究目的:本文通过案例研究形式开展相关评估,针对不同机器学习分类器进行对比分析,重点围绕数据清洗与预处理、特征选择、参数设置及模型选型等核心环节展开。 研究方法:本实验案例研究采用困境分析访谈语料库-奥兹巫师(Distress Analysis Interview Corpus - Wizard-of-Oz, DAIC-WOZ)数据集,该数据集专为抑郁症、焦虑症及创伤后应激障碍(PTSD)等精神疾病的诊断研究设计。 主要研究结果:除完成各类替代技术的评估外,本研究构建的随机森林(Random Forest)与XGBoost模型准确率可达84%左右,显著优于现有可比文献中支持向量机(SVM)模型72%的准确率水平。 研究结论:针对抑郁症检测任务的机器学习分类算法与自然语言处理技术开展更全面的评估,能够从优化实验设置与提升模型性能两方面推动该领域的研究前沿发展。



