ASPIR-CLIN Dataset
收藏资源简介:
Overview The ASPIR-CLIN Dataset is a clinical dataset designed to support research involving aspiration pneumonia, dysphagia assessment, swallowing disorders, and associated clinical risk factors. The dataset contains structured clinical information collected from patients evaluated through multiple clinical instruments and screening procedures, including swallowing assessment scales and hospitalization-related variables. The primary goal of this dataset is to facilitate: * Clinical research involving aspiration pneumonia;* Dysphagia risk analysis;* Statistical analyses of swallowing-related conditions;* Machine learning applications in healthcare;* Development of predictive clinical models. Dataset Contents The repository currently contains the following files: aspir_clin_dataset.csv = Main clinical datasetaspir_clin_data_dictionary.csv = Variable descriptions and coding informationaspir_clin_loader.py = Scikit-learn-compatible Python dataset loaderREADME.md = Dataset documentation Clinical Variables The dataset includes variables associated with: * Demographic information;* Feeding route;* Swallowing assessments;* Neurological conditions;* Respiratory conditions;* Cardiac conditions;* Nutritional indicators;* Functional limitations;* ICU and hospitalization information;* Aspiration pneumonia risk factors. Examples of included clinical instruments: * FEES;* FOIS;* EAT-10;* Glasgow Coma Scale. Data Collection Institution * Tuiuti University of Paraná Data Collection Period The sample was collected between May 2024 and March 2025 through a retrospective analysis of electronic medical records referring to the years 2022 and 2023. The data were extracted from the Tasy electronic system used by Hospital Nossa Senhora das Graças (HNSG), located in Curitiba, Paraná, Brazil, a reference institution for high-complexity clinical and surgical treatments. Population Hospitalized adults. Inclusion Criteria The study included patients of both sexes, aged 18 years or older, who underwent the institutional protocol for the investigation of bronchoaspiration and who underwent a Fiberoptic Endoscopic Evaluation of Swallowing (FEES) during hospitalization and/or a speech-language pathology clinical assessment. Exclusion Criteria The exclusion criteria included the absence of a speech-language pathology assessment record in the electronic medical record, incomplete data regarding the established variables, and age under 18 years. Ethical Considerations All patient information was anonymized prior to dataset publication. This study was conducted according to institutional ethical guidelines. Ethics Committee Approval Approved by the Research Ethics Committee under protocol number 75183123.9.3001.0269 Anonymization The following anonymization procedures were applied: * Removal of direct patient identifiers;* Replacement of medical record identifiers by synthetic patient IDs (`P001`, `P002`, etc.);* Removal of personally identifiable information. Dataset Structure Main Table Format * One row corresponds to one patient;* One column corresponds to one clinical variable. Python Loader The repository includes a Python loader (`aspir_clin_loader.py`) designed to provide a scikit-learn-compatible interface for loading the ASPIR-CLIN Dataset. The loader returns a `Bunch` object similar to datasets available in the scikit-learn library (e.g., the Iris dataset), providing direct access to: * `data`: feature matrix;* `target`: target variable;* `feature_names`: predictor variable names;* `target_names`: target class names;* `frame`: complete pandas DataFrame;* `DESCR`: dataset description. Example from aspir_clin_loader import load_aspir_clin dataset = load_aspir_clin("aspir_clin_dataset.csv") print(dataset.data.shape) print(dataset.target.shape) print(dataset.feature_names) print(dataset.target_names) print(dataset.DESCR) By default, the loader uses `penetration_risk` as the target variable, making the dataset readily applicable to machine learning, statistical analysis, and clinical decision support research. Missing Values No missing values were identified in Version 1.0 of the dataset. Future versions may include missing values represented as NA or empty cells. Variable Coding Detailed variable descriptions and codifications are available in: aspir_clin_data_dictionary.csv Potential Applications The ASPIR-CLIN Dataset may support studies involving: * Aspiration pneumonia prediction;* Dysphagia severity analysis;* Statistical analysis of clinical risk factors;* Explainable artificial intelligence in healthcare;* Clinical decision support systems;* Machine learning and data mining applications. Funding Coordination for the Improvement of Higher Education Personnel (CAPES) was responsible for granting the study scholarship to the lead author. Contact Information Rosane Sampaio Tuiuti University of Paranárosanesampaio.fono@gmail.com Recommended Citation If you use the ASPIR-CLIN Dataset in your research, please cite it as: Ramos, M. E. M., Tomasi Junior, D. L., & Sampaio, R. (2026). ASPIR-CLIN Dataset: Aspiration Pneumonia Clinical Dataset (Version 1.0) [Data set]. Zenodo. https://10.5281/zenodo.20527814 Dataset Authors - Maria Eduarda Medeiros Ramos- Darci Luiz Tomasi Junior- Rosane Sampaio Acknowledgments We would like to thank the Health Artificial Intelligence Center (NIAS) for the support in the development of this study, and the Coordination for the Improvement of Higher Education Personnel (CAPES) for granting a study scholarship. License This dataset is distributed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). https://creativecommons.org/licenses/by/4.0/ Disclaimer This dataset is intended exclusively for scientific and research purposes. The authors are not responsible for clinical decisions derived from the use of this dataset.
概述 本数据集为ASPIR-CLIN临床数据集,旨在支持吸入性肺炎、吞咽困难评估、吞咽障碍及相关临床危险因素相关研究。 该数据集包含通过多种临床工具与筛查流程(含吞咽评估量表及住院相关变量)收集自患者的结构化临床信息。 本数据集的核心目标是推动以下研究: * 吸入性肺炎相关临床研究 * 吞咽困难风险分析 * 吞咽相关病症的统计分析 * 医疗领域机器学习应用 * 临床预测模型开发 ## 数据集内容 当前仓库包含以下文件: * aspir_clin_dataset.csv:主临床数据集 * aspir_clin_data_dictionary.csv:变量说明与编码信息 * aspir_clin_loader.py:兼容scikit-learn的Python数据集加载器 * README.md:数据集文档 ## 临床变量 数据集涵盖以下类别相关变量: * 人口统计学信息 * 喂养途径 * 吞咽评估结果 * 神经系统病症 * 呼吸系统病症 * 心血管系统病症 * 营养指标 * 功能受限情况 * 重症监护与住院相关信息 * 吸入性肺炎危险因素 本数据集包含的临床工具示例如下: * 纤维内镜吞咽检查(FEES, Fiberoptic Endoscopic Evaluation of Swallowing) * 功能口腔进食量表(FOIS, Functional Oral Intake Scale) * 进食评估问卷-10(EAT-10, Eating Assessment Tool-10) * 格拉斯哥昏迷量表(Glasgow Coma Scale) ## 数据采集 ### 采集机构 巴拉那州图伊图大学(Tuiuti University of Paraná) ### 数据采集周期 样本采集于2024年5月至2025年3月期间,通过回顾性分析2022至2023年的电子病历完成。数据提取自位于巴西巴拉那州库里提巴的诺萨森霍拉达斯格拉卡斯医院(Hospital Nossa Senhora das Graças, HNSG)的Tasy电子系统,该医院是高复杂度临床与外科治疗的标杆医疗机构。 ### 研究人群 住院成年患者。 ### 纳入标准 本研究纳入年满18周岁、性别不限的患者,这些患者接受了机构制定的支气管吸入调查方案,并在住院期间接受了纤维内镜吞咽检查(FEES)或语言病理学临床评估。 ### 排除标准 排除标准包括电子病历中无语言病理学评估记录、既定变量数据不全,以及年龄未满18周岁的患者。 ## 伦理考量 所有患者信息在数据集发布前均已完成匿名化处理。本研究严格遵循机构伦理指南开展。 ### 伦理委员会审批 本研究经研究伦理委员会批准,审批编号为75183123.9.3001.0269。 ### 匿名化处理 采用以下匿名化流程: * 移除直接患者标识符 * 将病历标识符替换为合成患者ID(如`P001`、`P002`等) * 移除所有个人可识别信息 ## 数据集结构 ### 主表格式 * 每一行对应一名患者 * 每一列对应一项临床变量 ### Python加载器 仓库包含一个Python加载器(`aspir_clin_loader.py`),旨在提供兼容scikit-learn的接口以加载ASPIR-CLIN数据集。该加载器返回与scikit-learn库中现有数据集(如鸢尾花数据集)类似的`Bunch`对象,可直接访问以下属性: * `data`:特征矩阵 * `target`:目标变量 * `feature_names`:预测变量名称 * `target_names`:目标类别名称 * `frame`:完整的Pandas DataFrame * `DESCR`:数据集说明文档 ### 使用示例 python from aspir_clin_loader import load_aspir_clin dataset = load_aspir_clin("aspir_clin_dataset.csv") print(dataset.data.shape) print(dataset.target.shape) print(dataset.feature_names) print(dataset.target_names) print(dataset.DESCR) 默认情况下,加载器以`penetration_risk`作为目标变量,使该数据集可直接应用于机器学习、统计分析及临床决策支持相关研究。 ## 缺失值情况 该数据集1.0版本未发现缺失值。未来版本可能会使用NA或空单元格表示缺失值。 ## 变量编码 详细的变量说明与编码规则可参见`aspir_clin_data_dictionary.csv`文件。 ## 潜在应用场景 ASPIR-CLIN数据集可支持以下相关研究: * 吸入性肺炎预测 * 吞咽困难严重程度分析 * 临床危险因素统计分析 * 医疗领域可解释人工智能 * 临床决策支持系统 * 机器学习与数据挖掘应用 ## 资助情况 高等教育人才培养提升协调委员会(CAPES, Coordenação de Aperfeiçoamento de Pessoal de Nível Superior)为第一作者提供了研究奖学金。 ## 联系方式 Rosane Sampaio 巴拉那州图伊图大学 rosanesampaio.fono@gmail.com ## 推荐引用格式 若您在研究中使用ASPIR-CLIN数据集,请按以下格式引用: Ramos, M. E. M., Tomasi Junior, D. L., & Sampaio, R. (2026). ASPIR-CLIN Dataset: Aspiration Pneumonia Clinical Dataset (Version 1.0) [数据集]. Zenodo. https://10.5281/zenodo.20527814 ## 数据集作者 - Maria Eduarda Medeiros Ramos - Darci Luiz Tomasi Junior - Rosane Sampaio ## 致谢 感谢健康人工智能中心(NIAS, Núcleo de Inteligência Artificial em Saúde)为本研究开发提供的支持,以及高等教育人才培养提升协调委员会(CAPES)提供的研究奖学金。 ## 授权协议 本数据集采用知识共享署名4.0国际许可协议(CC BY 4.0)进行分发。 https://creativecommons.org/licenses/by/4.0/ ## 免责声明 本数据集仅用于科学研究目的。作者不对基于本数据集使用所产生的临床决策承担任何责任。



