An open-access EEG dataset for speech decoding
收藏资源简介:
本数据集为公开可用的EEG数据集,用于语音解码研究,探索发音和协同发音的作用。数据集包含两个验证集(N=8和N=16),用于音素和单词级别的分类,以及音素的发音特性。EEG信号通过64个通道记录,受试者在听和重复六个辅音和五个元音时进行。通过在不同的语音环境中组合单个音素,产生了四十个辅音-元音对、二十个真实单词和二十个伪词的协同发音变异。在控制条件和经颅磁刺激下,针对特定的发音过程抑制或增强EEG信号。
This dataset is a publicly available EEG dataset designed for speech decoding research, exploring the roles of articulation and coarticulation. The dataset includes two validation sets (N=8 and N=16) for phoneme and word-level classification, as well as the articulatory characteristics of phonemes. EEG signals were recorded through 64 channels while subjects listened to and repeated six consonants and five vowels. By combining individual phonemes in various speech contexts, forty consonant-vowel pairs, twenty real words, and twenty pseudowords with coarticulatory variations were generated. Under controlled conditions and transcranial magnetic stimulation, EEG signals were either suppressed or enhanced for specific articulatory processes.
数据集概述
数据集名称
An open-access EEG dataset for speech decoding: Exploring the role of articulation and coarticulation
数据集内容
- EEG信号记录:从64个通道记录EEG信号。
- 实验任务:受试者听并重复六个辅音和五个元音。
- 数据组成:
- 四十个辅音-元音对,包含共通音变化。
- 二十个真实单词。
- 二十个伪词。
- 控制条件下的数据和经颅磁刺激(TMS)下的数据,用于抑制或增强与特定发音过程相关的EEG信号。
数据集用途
- 用于语音解码的脑机接口(BCI)系统的研究。
- 支持在音素和单词级别的分类,以及根据音素的发音特性进行分类。
数据集规模
- 两个验证数据集:N=8和N=16。
代码和数据可用性
- 数据和代码可在OSF和GitHub上获取,遵循CC BY 4.0许可。
- 代码位于
Study/EEG_Data_Processing/Code文件夹,用于技术验证部分的分析。 - 信号处理技术的结果也存储在同一文件夹中。
数据处理方法
- 使用ICA和信号清洗技术,基于EEGLab库(MATLAB 2022.0和2022.1版本)进行数据处理。
图示说明
- Figure 1:展示了数据处理的代码结构。




