遇见数据集

Synthetic data using TVAE.

收藏
Figshare2025-06-02 更新2026-04-28 收录
官方服务:

资源简介:

Occupational stress is a major concern for employers and organizations as it compromises decision-making and overall safety of workers. Studies indicate that work-stress contributes to severe mental strain, increased accident rates, and in extreme cases, even suicides. This study aims to enhance early detection of occupational stress through machine learning (ML) methods, providing stakeholders with better insights into the underlying causes of stress to improve occupational safety. Utilizing a newly published workplace survey dataset, we developed a novel feature selection pipeline identifying 39 key indicators of work-stress. An ensemble of three ML models achieved a state-of-the-art accuracy of 90.32%, surpassing existing studies. The framework’s generalizability was confirmed through a three-step validation technique: holdout-validation, 10-fold cross-validation, and external-validation with synthetic data generation, achieving an accuracy of 89% on unseen data. We also introduced a 1D-CNN to enable hierarchical and temporal learning from the data. Additionally, we created an algorithm to convert tabular data into texts with 100% information retention, facilitating domain analysis with large language models, revealing that occupational stress is more closely related to the biomedical domain than clinical or generalist domains. Ablation studies reinforced our feature selection pipeline, and revealed sociodemographic features as the most important. Explainable AI techniques identified excessive workload and ambiguity (27%), poor communication (17%), and a positive work environment (16%) as key stress factors. Unlike previous studies relying on clinical settings or biomarkers, our approach streamlines stress detection from simple survey questions, offering a real-time, deployable tool for periodic stress assessment in workplaces.

职业压力是雇主与各类组织面临的核心关切之一,因其会损害员工的决策能力与整体安全状况。已有研究表明,工作压力会引发严重的精神负荷、事故率上升,极端情况下甚至会导致自杀行为。本研究旨在通过机器学习(Machine Learning, ML)方法实现职业压力的早期检测,为利益相关方提供压力根源的深度洞察,以优化职业安全保障水平。本研究利用最新公开的职场调研数据集,构建了全新的特征选择流程,筛选出39项工作压力关键表征指标。通过集成三种机器学习模型,本研究实现了90.32%的当前最优准确率,优于现有同类研究成果。本框架的泛化性能通过三步验证方法得到充分验证:留出验证(holdout-validation)、10折交叉验证(10-fold cross-validation)以及基于合成数据生成的外部验证(external-validation),在未知测试集上的准确率达到89%。此外,本研究引入了一维卷积神经网络(1D-CNN),以实现对数据的层级化与时序化学习。我们还开发了一种可100%保留原始信息的表格数据转文本算法,可借助大语言模型(Large Language Model, LLM)开展领域分析,结果显示职业压力与生物医学领域的关联度高于临床领域与通用领域。消融实验(ablation study)验证了所提出的特征选择流程的有效性,并揭示出社会人口学特征是影响压力的最重要因素。可解释AI(Explainable AI, XAI)技术识别出三大核心压力因素:过度工作负荷与权责模糊(占比27%)、沟通不畅(占比17%)以及工作环境正向性不足(占比16%)。与以往依赖临床场景或生物标志物的研究不同,本研究仅需通过简易调研问卷即可实现压力检测,流程更为简洁高效,可提供可实时部署的工具,用于职场定期压力评估。

创建时间:
2025-06-02
二维码
社区交流群
二维码
科研交流群
商业服务