遇见数据集

Multimodal Spirometry Dataset for Quality Assessment: Tabular Parameters and Flow-Volume Loop Images

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This dataset accompanies the study titled "Explainable Machine Learning for Quality Assessment of Spirometry Tests: A Multimodal Framework." It comprises 2,768 spirometry records collected retrospectively from a tertiary care hospital (Yozgat Bozok University, Turkey) between January 2020 and October 2024. Dataset Structure The dataset consists of two components linked by a unique patient identifier (Id): spirometry_tabular_data.csv — A structured tabular file containing 2,768 records with 44 columns. The first column (Id) serves as the unique patient identifier that maps each row to its corresponding flow-volume loop image. The remaining columns include demographic parameters (Gender, Age, Height, Weight), environmental variables (Temperature, Barometric Pressure), spirometric measurements (FEV1, FVC, FEF25-75%, PEF, FET100%, FIVC, FIV1, FEF/FIF50, Vol Extrap) with their PRED, BEST, and %PRED values, and a binary target label (class: 1 = adequate, 0 = inadequate). flow_volume_loop_images/ — A folder containing 2,768 PNG images of flow-volume loops. Each image is named as {Id}.png (e.g., 1.png, 100.png, 2336.png), directly corresponding to the Id column in the tabular file. Linkage: The Id column in the CSV file and the image filenames establish a one-to-one mapping between tabular records and flow-volume loop images. Annotations: The acceptability of spirometric maneuvers (adequate/inadequate) was labeled by expert physicians according to ATS/ERS 2005 guidelines and updated technical recommendations by Stanojevic et al. Class Distribution: Approximately balanced — 52.9% adequate (class 1, n=1,464) and 47.1% inadequate (class 0, n=1,304). Missing Data: The data quality score is 70.6%, with notable missing rates in certain features (e.g., FEV1, FVC, and FEF values exceeding 90% missingness in some subsets). Missing values are encoded as "Na" in the CSV file. Usage: This dataset was used to train and evaluate multiple classification models including tree-based ensemble methods for tabular data, an EfficientNet-B2-based deep learning model for image data, and a late fusion architecture integrating both modalities. Explainability analyses were conducted using SHAP, LIME, and Grad-CAM. Ethical Approval: Protocol No: 2024-GOKAEK-248_18.09.2024_148, Ethics Committee of Yozgat Bozok University. Associated Article: This dataset is associated with a manuscript currently under review. The full citation and DOI of the related article will be added upon publication.

本数据集配套于题为《可解释机器学习用于肺量测定检查质量评估:一种多模态框架》的研究。数据集包含2768条肺量测定记录,于2020年1月至2024年10月间从土耳其约兹加特博佐克大学附属三级医院回顾性采集获得。 数据集结构:本数据集由两个组件构成,二者通过唯一患者标识符(Id)进行关联: 1. spirometry_tabular_data.csv:一份结构化表格文件,包含2768条记录与44列数据。首列(Id)作为唯一患者标识符,将每一行与对应的流量-容积环图像进行绑定映射。其余列涵盖人口统计学参数(性别、年龄、身高、体重)、环境变量(温度、大气压)、肺量测定指标(FEV1、FVC、FEF25-75%、PEF、FET100%、FIVC、FIV1、FEF/FIF50、Vol Extrap)及其PRED、BEST与%PRED值,同时包含二分类目标标签(类别:1代表合格,0代表不合格)。 2. flow_volume_loop_images/:一个存储2768张流量-容积环PNG图像的文件夹。每张图像以{Id}.png格式命名(例如1.png、100.png、2336.png),与表格文件中的Id列一一对应。 关联方式:CSV文件中的Id列与图像文件名,实现了表格记录与流量-容积环图像的一一映射。 标注说明:肺量测定操作的可接受性(合格/不合格)由专科医师依据美国胸科学会/欧洲呼吸学会(ATS/ERS)2005年指南,以及Stanojevic等人更新的技术建议完成标注。 类别分布:数据集类别近似平衡——合格样本(类别1,n=1464)占比52.9%,不合格样本(类别0,n=1304)占比47.1%。 缺失数据:本数据集的数据质量得分为70.6%,部分特征存在显著缺失率(例如部分子集的FEV1、FVC及FEF指标缺失率超过90%)。CSV文件中缺失值以"Na"进行编码。 数据集用途:本数据集被用于训练与评估多种分类模型,包括面向表格数据的树集成方法、面向图像数据的EfficientNet-B2深度学习模型,以及融合两种模态的晚融合架构。同时采用SHAP、LIME及Grad-CAM开展可解释性分析。 伦理审批:本研究已通过约兹加特博佐克大学伦理委员会审批,审批编号为2024-GOKAEK-248_18.09.2024_148。 关联文章:本数据集关联的手稿目前处于审稿阶段。相关文章的完整引用信息及DOI将在正式发表后补充。

创建时间:
2026-03-11
二维码
社区交流群
二维码
科研交流群
商业服务