ScientISST MOVE: Annotated Multimodal Naturalistic Dataset Recorded During Everyday Life Activities Using Wearable Devices
收藏资源简介:
A multi-modality, multi-activity, and multi-subject dataset of wearable biosignals. Modalities: ECG, EMG, EDA, PPG, ACC, TEMP Main Activities: Lift object, Greet people, Gesticulate while talking, Jumping, Walking, and Running Cohort: 17 subjects (10 male, 7 female); median age: 24 Devices: 2x ScientISST Core + 1x Empatica E4 Body Locations: Chest, Abdomen, Left bicep, wrist and index finger No filter has been applied to the signals, but the correct transfer functions were applied, so the data is given in relevant unis (mV, uS, g, ºC). ======== There are three formats available: a) LTBio's Biosignal files. Should be open like: x = Biosignal.load(path) LTBio Package: https://pypi.org/project/LongTermBiosignals/ Under the directory biosignal, the following tree structure is found: subject/x.biosignal, where subject is the subject’s code, and x is any of the following { acc_chest, acc_wrist, ecg, eda, emg, ppg, temp }. Each file includes the signals recorded from every sensor that acquires the modality after which the file is named, independently of the device. Channels, activities and time intervals can be easily indexed with the index operator []: https://ltbio.readthedocs.io/en/latest/learn/basic/ltbio101.html A sneak peak of the signals can also be quickly plotted with: x.preview.plot() Any Biosignal can be easily converted to NumPy arrays or DataFrames, if needed. b) CSV files. Can be open like: x = pandas.read_csv(path) Pandas Package: https://pypi.org/project/pandas/ These files can be found under the directory csv, named as subject.csv, where subject is the subject’s code. There is only one file per subject, containing their full session and all biosignal modalities. When read as tables, the time axis is in the first column, each sensor is in one of the middle columns, and the activity labels are in the last column. In each row are the samples of each sensor, if any, at each timestamp. At any given timestamp, if there is no sample for a sensor, it means the acquisition was interrupted for that sensor, which happens between activities, and sometimes for short periods during the running activity. Also in each row, on the last column, is one or more activity labels, if an activity was taking place at that timestamp. If there are multiple annotations, the labels are separated by commas (e.g 'run,sprint'). If there are no annotations, the column is empty for that timestamp. In order to provide a tabular format with sensors with different sampling frequencies, the sensors with sampling frequency lower than 500 Hz were upsampled to 500 Hz. This way, the tables are regularly sampled, i.e., there is a row every 2 ms. If a sensor was not acquiring at a given timestamp, the corresponding cell with be empty. So, not only the segments with samples are regularly sampled, but the interruptions are also discretised. This means that if, after an interruption, a sensor starts acquiring at a non regular timestamp, the first sample will be written on the previous or the following timestamp, by half-up rounding. Naturally, this process cumulatively introduces lags in the table, some of which cancel out. Each individual lag is no longer than half the sampling period (1 ms), hence negligible. The cumulative lags are no longer than 200 ms for all subjects, which is also negligible. Nevertheless, only the LBio's Biosignal format preserves the exact original timestamps (10E-6 precision) of all samples and the original sampling frequencies. ================ Both include annotations of the activities, however LTBio bio signal files have better time resolution and include clinical data and demographic data as well. c) EDF+ files. Can be open like: x = mne.io.read_raw_edf(path) MNE Package: https://mne.tools/stable/index.html Under the directory edf, the following tree structure is found: subject/x.edf, where subject is the subject’s code, and x is any of the following { empathic, scientisst_chest, scientisst_forearm }. Each file includes the signals recorded from every device after which the file is named, independently of the modality. Notes: Original sampling frequencies are maintained. Original units are maintained. Signal is NaN during recording interruptions. Events are in EDF annotations. Biosignal and patient notes are not maintained. The signals can be quickly plotted with: x.plot(). Make sure you have interactive Matplotlib activated. At first, you might have to decrease the scaling in order to correctly inspect them. tip: use the minus (-) kay in your keyboard as many times as necessary to reduce the scaling.
一款面向可穿戴生物信号的多模态、多活动场景、多受试者数据集。采集模态:心电图(ECG)、肌电图(EMG)、皮肤电反应(EDA)、光电容积描记图(PPG)、加速度计信号(ACC)、体温信号(TEMP)。主要活动场景:举取物品、问候他人、交谈时手势动作、跳跃、行走与跑步。受试者队列:共17名受试者(男性10名,女性7名),年龄中位数为24岁。采集设备:2台ScientISST Core设备 + 1台Empatica E4设备。传感器佩戴位置:胸部、腹部、左二头肌、手腕与食指。本数据集未对信号施加滤波处理,但已应用正确的传递函数,因此数据采用对应物理单位:毫伏(mV)、微西门子(uS)、克(g)与摄氏度(ºC)。 ======= 本数据集提供三种存储格式: a) LTBio生物信号文件格式:可通过"x = Biosignal.load(path)"加载。配套LTBio工具包下载地址:https://pypi.org/project/LongTermBiosignals/。该格式文件存储于"biosignal"目录下,目录结构为"subject/x.biosignal",其中"subject"为受试者编码,"x"为以下任一模态:acc_chest、acc_wrist、ecg、eda、emg、ppg、temp。每个文件包含对应命名模态下所有传感器采集的信号,与具体设备无关。可通过索引运算符[]快速对通道、活动与时间区间进行索引,具体操作可参考官方文档:https://ltbio.readthedocs.io/en/latest/learn/basic/ltbio101.html。也可通过"x.preview.plot()"快速预览信号波形。若有需要,可将任意生物信号文件轻松转换为NumPy数组或Pandas数据框(DataFrames)。 b) CSV文件格式:可通过"x = pandas.read_csv(path)"加载,配套Pandas工具包下载地址:https://pypi.org/project/pandas/。该格式文件存储于"csv"目录下,命名为"subject.csv",其中"subject"为受试者编码。每名受试者仅对应一个文件,包含其完整采集会话与所有生物信号模态的数据。以表格形式读取后,第一列为时间轴,中间各列为各传感器采集的数据,最后一列为活动标签。每一行对应一个时间戳下的各传感器采样值(若存在)。若某时间戳下某传感器无采样数据,则表示该传感器此时采集中断——中断通常发生在活动切换期间,有时也会在跑步活动中短暂出现。每一行的最后一列包含对应时间戳下正在进行的一项或多项活动标签,多标签间以逗号分隔(例如"run,sprint");若无活动标签,则该单元格为空。为适配不同采样频率的传感器以表格形式存储,本数据集将采样频率低于500Hz的传感器数据上采样至500Hz,因此表格为均匀采样格式,即每2ms生成一行数据。若某传感器在某时间戳未采集数据,则对应单元格为空。如此一来,不仅存在采样数据的片段为均匀采样,采集中断也被离散化处理:若传感器在中断后以非均匀时间戳恢复采集,则其首个采样值将通过半向上取整规则匹配至前一个或后一个时间戳。该处理会累计引入时间偏移,部分偏移可相互抵消,单个偏移时长不超过采样周期的一半(1ms),因此可忽略不计;所有受试者的累计偏移时长均不超过200ms,同样可忽略不计。仅LTBio生物信号格式保留了所有采样点的精确原始时间戳(精度为10^-6秒)与原始采样频率。 =============== 两种CSV与LTBio生物信号文件均包含活动标注,但LTBio生物信号文件具备更优的时间分辨率,同时还包含临床数据与人口统计学信息。 c) EDF+文件格式:可通过"x = mne.io.read_raw_edf(path)"加载,配套MNE工具包官网地址:https://mne.tools/stable/index.html。该格式文件存储于"edf"目录下,目录结构为"subject/x.edf",其中"subject"为受试者编码,"x"为以下任一设备组:empathic、scientisst_chest、scientisst_forearm。每个文件包含对应命名设备组下所有传感器采集的信号,与具体模态无关。注意事项:保留原始采样频率与物理单位;采集中断期间的信号值为NaN;事件信息存储于EDF注释中;未保留生物信号与患者备注信息。可通过"x.plot()"快速绘制信号波形,需确保已启用交互式Matplotlib。初次查看时可能需要调整缩放比例以正确观测波形,提示:可根据需要多次按下键盘上的减号(-)键以降低缩放级别。




