遇见数据集

The SpeakUnique Corpus of Heard and Imagined Natural Speech (CHINS)

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

Overview The Corpus of Heard and Imagined Natural Speech (CHINS) is a single-participant electroencephalography (EEG) corpus comprising over 22 hours of time-aligned natural speech data, with approximately 11 hours of heard speech and 11 hours of imagined speech. The corpus was recorded from a single 26-year-old right-handed male participant across nine recording sessions using a BioSemi ActiveTwo 64-channel EEG system at 1,024 Hz. The dataset contains 73 recordings of approximately 20 minutes each. The corpus was designed to enable direct comparison of neural responses to heard and imagined speech under closely matched temporal conditions. The audiovisual stimuli comprise public-domain recordings of The Adventures of Sherlock Holmes, accompanied by word-by-word visual prompts. Following presentation of the heard speech, the participant was instructed to imagine the same speech with the same rhythm and prosody. Automatic forced alignment provides phone-level temporal annotations for both the heard and imagined phases. Contents The dataset is organised according to the Brain Imaging Data Structure (BIDS), version 1.10.0. It includes raw EEG recordings and BIDS sidecars, electrode positions, audiovisual stimuli, forced-alignment files, and detailed event annotations. Events are labelled at the phone level and distinguish heard (listen) from imagined (think) speech, with within-word phone position also encoded. The recordings additionally contain synchronisation beeps, whose events are documented in the dataset and can be used to identify and exclude periods of EEG contamination surrounding the evoked response. Processing The distribution includes the complete processing pipeline used to clean, epoch, and average the EEG data into event-related potentials (ERPs). The pipeline is implemented using MNE-Python and produces cleaned continuous data, event-locked epochs, and ERPs, including frequency-band-specific outputs. Because the corpus contains approximately 469,000 events, processing the complete dataset requires ample computational resources; detailed processing requirements and instructions are provided in the accompanying README. Referencing CHINS is intended to support research on the neural representation of natural speech, with particular emphasis on the relationship between auditory speech perception and imagined speech (inner speech). It accompanies its associated 2026 Interspeech publication. If you use CHINS, reference both: S. Wellington, O. Watts, D. Coyle, and B. Metcalfe, “Shared phone-level neural representations of auditory perception and ‘inner voice’ production: One-to-one mapping using a single-subject EEG corpus of heard and imagined natural speech,” in Proc. Interspeech 2026, 2026, pp. 4864–4869, doi: 10.21437/Interspeech.2026-2683. S. Wellington, O. Watts, D. Coyle, and B. Metcalfe, The Corpus of Heard and Imagined Natural Speech (CHINS), Zenodo, Sep. 26, 2026, doi: 10.5281/zenodo.22974624. Legal The dataset is released under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) licence. It is made available for non-commercial purposes, including academic research and teaching, with appropriate attribution. Users may process the data for their own research; the licensing conditions governing redistribution and commercial use are described in the accompanying LICENSE file. Access Access available on request: please contact info@speakunique.co.uk with details of your affiliation and description of how you will use the data to request access and download permissions.Full documentation of the dataset structure, event annotations, processing pipeline, computational requirements, references, ethics approval, and licensing is provided in the accompanying README.CHINS is research from SpeakUnique: https://www.speakunique.co.uk/research/CHINS.

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务