遇见数据集

Dataset for Analyzing Psychological Barriers to Smoking Cessation Using Machine Learning in Bangladesh

收藏
Mendeley Data2026-08-04 收录
官方服务:

资源简介:

This dataset accompanies the research "Analyzing Psychological Barriers to Smoking Cessation Using Machine Learning" and contains anonymized survey data collected from current and former smokers in Bangladesh. The study was conducted as part of a Bachelor of Science Final Year Design Project at the Department of Computer Science and Engineering, Daffodil International University. Data were collected using a structured 21-item questionnaire administered through both Google Forms and face-to-face interviews to increase participation from individuals with limited access to digital platforms. The questionnaire was independently reviewed by two medical experts using the Content Validity Index (CVI) prior to data collection. A total of 901 responses were collected. After removing records that were not eligible for predictive modelling, 673 model-ready responses were retained for analysis. The dataset contains demographic information, smoking history, previous quit attempts, and ten psychological variables measured on five-point Likert scales. The target variable indicates whether a participant relapsed within one month after a previous quit attempt. The dataset was developed to support research on smoking cessation, public health informatics, explainable artificial intelligence (XAI), and machine learning. It can be used for binary classification, feature selection, class imbalance research, predictive modelling, explainability analysis using SHAP, and benchmarking machine learning algorithms. All personally identifiable information has been removed before publication to protect participant privacy.

本数据集配套支持研究论文《使用机器学习分析戒烟的心理障碍》("Analyzing Psychological Barriers to Smoking Cessation Using Machine Learning"),包含从孟加拉国当前吸烟者与既往吸烟者中收集的匿名调研数据。本研究作为孟加拉国达芙妮国际大学(Daffodil International University)计算机科学与工程系的理学学士毕业设计项目开展。 数据收集采用包含21个条目的结构化问卷,通过谷歌表单(Google Forms)与面对面访谈两种方式进行,以提升数字平台使用受限人群的参与率。在数据收集前,两名医学专家已通过内容效度指数(Content Validity Index, CVI)对问卷进行独立审核。 本次调研共回收901份问卷。剔除不适用于预测建模的记录后,最终保留673份可用于建模的有效样本供分析使用。本数据集包含人口统计学信息、吸烟史、既往戒烟尝试情况,以及10项基于5级李克特量表(five-point Likert scales)测量的心理变量。目标变量用于指示参与者在既往戒烟尝试后1个月内是否复吸。 本数据集旨在支持戒烟、公共卫生信息学、可解释人工智能(Explainable Artificial Intelligence, XAI)以及机器学习领域的相关研究。其可应用于二分类任务、特征选择、类别不平衡研究、预测建模、基于SHAP的可解释性分析,以及机器学习算法基准测试。所有个人可识别信息均已在发布前移除,以保护受访者隐私。

创建时间:
2026-07-21
二维码
社区交流群
二维码
科研交流群
商业服务