Data for: The upper bound of the misclassification risk for the long-term myoelectric signal recognition based on the adaptive learning
收藏资源简介:
In order to provide the guideline for designing and evaluating the long-term myoelectric signal recognition methods based on the adaptive learning, we proposed theoretical models to describe how the boundary of the misclassification risk (BMR) change along parameters including, the adaptive learning times, the adaptive learning frequencies, the generalization ability of the predictive model, and the ratio of samples without supervised information during the adaptive learning. The models are built up based on the formulated adaptive learning process of the long-term myoelectric signal recognition, and the normalized definition of the concept drift and the concept sequence. Experiments based on both realistic long EMG data sequences (Realistic Concept-drift-rate Sequence, RCS) and simulated EMG data sequences with controllable permutations (Zero Concept-drift-rate Sequence, ZCS, and Constant Concept-drift-rate Sequence, CCS) are conducted to validate our theoretical analyses. All the data sequences are reorganized from the raw EMG dataset. The reorganization methods are presented in the paper. In the data set, we present the raw EMG dataset and the recognition result of the ZCSs, the CSSs and the RCSs. The folder “RawEMGData” contains the raw EMG dataset containing EMG data of 8 subjects noted as S1 to S8. For each subject, we acquire 24 data sessions, corresponding to the .dat files in the folder. The data sessions were acquired every a half hour. Each data session includes 4000 samples of 8 motion classes (500 samples for one motion class). Each row containing 8 data in the file represents a sample. The former 7 data in a row are the data acquired by 7 EMG channels, the last one data in a row is the motion label of the sample. The folder “CCSResult” contains the recognition result of the CCSs. The data are saved in the form of 5-dimention array in .mat files. The entry at position (i,j,k,l,m) saves the classification result of samples in lth session of kth session sequence, which are classified by the mth adaptive learner. The classification result is the number of samples with the ground truth label of j, and the classification result of i. The position varies from (1,1,1,1,1) to (8,8,100,L,8). The variable m=1 corresponds to the result of the IAL, m=2 to 8 correspond to the results of the RAL1 to RAL7. The maximum value of l is noted as L, which equals 8, 16, 18, 20, 22, 27, 32, 64, 128 for file “Bt_Seq1.mat” to file “Bt_Seq9.mat”.
为给基于自适应学习的长期肌电信号识别方法的设计与评估提供指导框架,我们构建了理论模型,用以描述误分类风险边界(Boundary of Misclassification Risk, BMR)随自适应学习轮次、自适应学习频率、预测模型泛化能力,以及自适应学习阶段无监督信息样本占比等参数的变化规律。该模型基于长期肌电信号识别的自适应学习过程形式化建模,并结合概念漂移(concept drift)与概念序列(concept sequence)的归一化定义构建。 我们通过两类实验验证理论分析结果:一是基于真实长期肌电数据序列(Realistic Concept-drift-rate Sequence, RCS),二是基于带有可控排列的模拟肌电数据序列(零概念漂移率序列(Zero Concept-drift-rate Sequence, ZCS)与恒定概念漂移率序列(Constant Concept-drift-rate Sequence, CCS))。所有数据序列均源自原始肌电数据集,具体重组方法已在论文中详述。 本数据集包含原始肌电数据以及ZCS、CCS与RCS的识别结果。文件夹「RawEMGData」内存储原始肌电数据集,涵盖8名受试者(标记为S1至S8)的肌电数据。每名受试者采集24个数据会话,对应文件夹内的.dat文件,数据会话每半小时采集一次。每个数据会话包含8种运动类别的4000个样本(每类运动对应500个样本)。文件中每行含8个数据,对应一个样本:前7个数据为7个肌电通道采集的信号,最后1个数据为该样本的运动类别标签。 文件夹「CCSResult」内存储CCS的识别结果,数据以5维数组形式保存于.mat文件中。数组位置(i,j,k,l,m)存储的分类结果为:使用第m个自适应学习器,对第k个会话序列的第l个会话中的样本进行分类后,真实标签为j、预测标签为i的样本数量。数组索引范围为(1,1,1,1,1)至(8,8,100,L,8)。其中m=1对应增量自适应学习器(IAL)的识别结果,m=2至8分别对应RAL1至RAL7的识别结果。l的最大值记为L,针对「Bt_Seq1.mat」至「Bt_Seq9.mat」这9个文件,L的取值分别为8、16、18、20、22、27、32、64、128。



