JapanEEG
收藏资源简介:
JapanEEG是由Araya公司创建的一个大规模多模态脑电数据集,专门针对日语语音生成研究。该数据集包含1020小时的同步记录头皮脑电图、面部肌电图和语音音频数据,来自三名健康日语母语者在开放词汇朗读任务中的多次会话,覆盖三种不同通道数(62-128通道)的脑电设备。数据采集过程采用纵向、多设备和多模态同步记录策略,通过文本游戏、书籍朗读和语音语料库等多种任务激发自然语音产出,并利用Silero VAD模型进行语音事件检测与标注。该数据集主要应用于语音解码、脑机接口、多模态信号处理以及跨会话和设备适应等领域,旨在解决语音相关脑电研究中数据规模有限、设备泛化能力不足以及肌肉伪影干扰等核心挑战。
JapanEEG is a large-scale multimodal electroencephalography (EEG) dataset developed by Araya Inc., exclusively tailored for Japanese speech generation research. It encompasses 1020 hours of synchronously recorded scalp EEG, facial electromyography (EMG), and speech audio data, sourced from multiple sessions of three healthy Japanese native speakers completing open-vocabulary oral reading tasks, across EEG devices with three distinct channel counts (62–128 channels). The data collection employs a longitudinal, multi-device, and multimodal synchronous recording paradigm, where natural speech production is elicited through diverse tasks including text games, book reading, and speech corpora. Speech event detection and annotation are implemented using the Silero VAD model. This dataset is primarily utilized in research areas such as speech decoding, brain-computer interfaces (BCIs), multimodal signal processing, as well as cross-session and cross-device adaptation. Its core objective is to address key challenges in speech-related EEG research, namely limited data scale, inadequate device generalization ability and muscle artifact interference.

- 1A 1000-hour EEG-EMG-audio dataset of Japanese speech productionAraya 公司 · 2026年



