FII BCI Corpus P300
收藏资源简介:
The FII BCI Corpus comprises a selection of BCI databases : P300 and Motor Imagery (MI), annotated and curated for research purposes in a collaborative project carried out at University Federico II of Naples and University Grenoble Alpes of Grenoble. Besides using this page and Zenodo API, the corpus can be easily installed using Eegle's download GUI — see Eegle.Database.downloadDB. The disk space requirements for the data here below is 14.2 GB. Along with EEG data and class labels, the corpus provides comprehensive metadata that allow to select the data for the study at hand and extract relevant information. This makes it particularly easy to carry out machine learning research on BCI data — see for example Tutorial ML 2. The data curation included: discarding EEG recordings flagged by the database authors as problematic holding corrupted data (e.g., NaN values during the trials) yielding numerical problems with standard data manipulations procedures offering close-to-chance performance in standard 2-class prediction tasks, conversion of data into µV with Float32 precision, class re-labeling using a standardized scheme, downsampling (if applicable) to ≤ 256 samples per second preventing aliasing, removal of non-EEG channels, such as EOG, EMG, reference, or ground electrodes, concatenation of runs from the same session with identical experimental condition, cleaning of NaN and zero values at the beginning and ending of the recordings conversion of the cleaned data from CSV to the NY format, easily accessible in any programming language. For the details on the discarding procedures for each file see the @discarded.md file for MI and P300. For all details on the construction of the corpus, see the FII BCI Corpus Overview.
FII BCI语料库(FII BCI Corpus)是由那不勒斯费德里科二世大学与格勒诺布尔阿尔卑斯大学合作开展的科研项目中,为研究目的标注并整理的精选脑机接口(BCI)数据库集合,涵盖P300与运动想象(Motor Imagery, MI)两类数据。 除可通过本页面与Zenodo应用程序编程接口(Application Programming Interface, API)获取外,该语料库还可通过Eegle的下载图形用户界面(Graphical User Interface, GUI)便捷安装——具体可参考Eegle.Database.downloadDB。 下述数据集的磁盘存储空间需求为14.2 GB。 该语料库除包含脑电图(Electroencephalogram, EEG)数据与类别标签外,还提供了全面的元数据,可用于筛选当前研究所需的数据并提取相关信息,极大简化了基于BCI数据开展机器学习研究的流程——例如可参考《Tutorial ML 2》。 数据整理工作包含以下内容: - 剔除经数据库作者标记为存在缺陷的EEG记录,具体涵盖:试验阶段出现非数值(NaN)值的数据损坏、标准数据处理流程触发数值异常、在标准二分类预测任务中性能接近随机猜测水平的记录; - 将数据转换为以微伏(µV)为单位的Float32精度格式; - 采用标准化方案重新标注类别标签; - (若适用)将采样率降至每秒≤256个样本,以避免混叠效应; - 移除非EEG通道,如眼电图(Electrooculogram, EOG)、肌电图(Electromyogram, EMG)、参考电极与接地电极; - 将同一实验条件下同一会话中的多次运行数据进行拼接; - 清理记录首尾的NaN与零值数据; - 将清理后的数据从逗号分隔值(Comma-Separated Values, CSV)格式转换为NY格式,该格式可在任意编程语言中便捷访问。 关于各文件的剔除细则,可参考运动想象(MI)与P300数据集对应的@discarded.md文件。 如需了解该语料库构建的全部细节,请参考《FII BCI Corpus Overview》。



