遇见数据集

MS-AASD: An EEG Dataset for Self-Initiated Auditory Attention Switch Decoding in Mixed Speech

收藏
Zenodo2026-01-23 更新2026-05-26 收录
官方服务:

资源简介:

To address limitations in prior auditory attention decoding (AAD) studies that rely on spatialized audio presentation and externally prompted attention switches, we introduce the mixed-speech auditory attention switch decoding dataset (MS-AASD). In many existing switching paradigms, attention changes are induced by explicit external manipulations, such as auditory or visual cues, predetermined switching times, or abrupt changes in spatial configuration (e.g., swapping speaker locations), which effectively impose switch timing by the experimental design. In contrast, MS-AASD records EEG under unspatialized mixed-speech conditions, removing directional cues between competing streams to reduce non-auditory confounds (e.g., eye movements), and captures self-initiated attention switches whose timing is not externally specified. Participants determine when to switch attention based solely on their own listening experience, without cues or predefined schedules. This design enables a controlled yet realistic setting to study endogenous attention dynamics under cue-free conditions, rather than responses to externally triggered events. Participants and Ethical Approval: Thirteen native Chinese adults (ages 21–28) with audiometrically confirmed normal hearing (pure-tone thresholds ≤ 25 dB HL at octave frequencies from 125 to 8000 Hz) participated in the study. All participants reported no history of neurological or psychiatric disorders. Written informed consent was obtained from all participants prior to the experiment. The study protocol was approved by the Institutional Ethical Review Board of Southern University of Science and Technology (No. 2022DZX003). Data Acquisition: EEG data were collected in a sound-attenuated booth while participants listened to auditory stimuli presented via headphones at a comfortable listening level (approximately 65 dB SPL). The experimental paradigm and stimulus presentation were controlled using E-Prime 3.0 software. EEG signals were recorded using a SynAmps RT system (NeuroScan) with 64 scalp electrodes (including the two mastoid reference electrodes) positioned according to the extended international 10–20 system, at a sampling rate of 500 Hz. Electrode impedances were maintained below 5 kΩ to ensure highquality signals. Throughout the experiment, EEG recordings, stimulus onset markers, and behavioral responses (keypress events) were acquired synchronously. Experimental Design and Procedure: Trials: Each participant completed 60 trials of 60 s each, organized into six blocks of ten trials with randomized presentation order. Short rest periods (5 s) were provided between trials, with longer self-paced breaks between blocks to mitigate fatigue. Task Instructions and Attention Switching: At the beginning of each trial, participants were instructed to select one of the two concurrent speech streams as the initial target during the first 2 s. Subsequently, they were asked to maintain continuous attention to one stream at a time and to perform a limited number of self-initiated attention switches (typically 2–5 per trial) without any constraints on when the switches should occur, based on their subjective listening experience (e.g., perceived engagement or listening effort). No external cues, prompts, or predefined schedules were provided to indicate switch timing. Importantly, participants were instructed to promptly indicate the currently attended speech stream by pressing the corresponding button (male-stream button vs. female-stream button) whenever they selected or changed their attentional target, including the initial target selection at the beginning of each trial. The instructed switching range served only to discourage extreme behaviors (e.g., no switching or random frequent switching), while the exact timing and motivation of each switch were determined endogenously by the participant rather than by the experimental protocol. Labeling via Keypress: Participants reported attention switches using two dedicated buttons, one mapped to the male speech stream and the other to the female speech stream. Thus, each keypress provides both a temporal marker of a self-initiated switch and an explicit indication of the newly attended target stream. The keypresses were used solely for reporting internally decided attention changes rather than inducing or guiding them, and single-finger input was adopted to minimize motor involvement. As internal attentional state changes cannot be directly observed, keypress timing provides an approximate and intentionally coarse alignment of switch events, reflecting realistic latency and uncertainty inherent in self-initiated paradigms. Training and Task Familiarization: Before the formal experiment, each participant completed two practice trials using the same task paradigm. These training trials were designed to ensure that participants clearly understood the concept of self-initiated attention switching, the absence of external switching cues, and the use of keypresses solely for reporting internally decided attention changes. Feedback was provided during the practice trials to confirm correct task understanding. Data Quality Control: To ensure data validity, strict quality control criteria were applied. Trials or participants were excluded if they did not comply with the task instructions, such as misunderstanding the self-initiated switching paradigm, producing excessively frequent or irregular keypresses indicative of random or inattentive behavior, or reporting difficulty maintaining attention during post-experiment debriefing. The instructed switching range (2–5 per trial) was used as a reference rather than a strict exclusion threshold. Only compliant data were retained in the released dataset. Window-Level Labeling and Temporal Alignment: In MS-AASD, the EEG signals and both speech streams (the attended target and the ignored distractor) are time-synchronized at the signal level. The only source of uncertainty lies in which speech stream is attended at a given time, rather than in the temporal alignment between modalities. Accordingly, switch labels are derived from keypress events and aligned to the EEG at the decision-window level rather than treated as precise switch points. EEG signals are segmented into fixed-length windows (e.g., 1 s), and for each window the attention label is assigned according to the dominant attentional state within that window, determined by the temporal overlap between the window and keypress-defined attention segments. This labeling strategy is consistent with standard auditory attention decoding practice, where decoding is performed at the window level rather than at exact behavioral event times. Although participants were instructed to press the button immediately upon deciding to switch attention, a variable delay between internal attentional changes and keypress events is unavoidable and is explicitly accommodated by the window-level labeling strategy. Majority-based window labeling therefore provides a natural and task-consistent representation of attentional state for streaming AAD systems. Switching Behavior Statistics: In total, 13 participants took part, yielding approximately 13 h of EEG recordings. The mean number of self-initiated attention switches per trial across participants was 3.37, with values ranging from approximately 1.1 to 6.6 switches per 60 s trial. Switching behavior varied across participants and across trials within the same participant, characterizing self-initiated switching patterns in MS-AASD. Intended Use and Research Applications: MS-AASD is designed to support research on dynamic and streaming auditory attention decoding under self-initiated switching, rather than precise switch-point detection. The dataset enables evaluation of EEG-based models that track the attended speech stream over time in the presence of endogenous attention changes and uncertain labels. It can be used to study (i) streaming versus offline AAD under attention switching, (ii) robustness of decoding models to label latency and ambiguity, (iii) temporal stabilization and history-aware decoding strategies, and (iv) EEG–speech matching or similarity-based decoding in cue-free mixed-speech settings. As such, MS-AASD complements existing sustained-attention or cue-driven AAD datasets by targeting a more realistic yet controlled attention-switching scenario. Dataset and Code Release: This initial release contains a randomly selected subset from four of the thirteen participants for research use only, and the full dataset will be made publicly available after the project’s final review in accordance with the deliverable requirements. This release includes: EEG responses to unspatialized two-speaker mixtures Clean speech from AISHELL (3 male, 3 female speakers), forming three male-female speaker pairs in total Per-trial mixed stimuli made by selecting non-overlapping male and female segments, RMS-normalizing, summing at 0 dB to a single-channel signal, and presenting the same signal to both ears (no spatial cues) EEG preprocessing code covering re-referencing, band-pass filtering, downsampling, and standardization

提供机构:
Zenodo
创建时间:
2025-09-18
二维码
社区交流群
二维码
科研交流群
商业服务