遇见数据集

A synchronized multimodal neuroimaging dataset to study brain language processing

收藏
OpenNeuro2022-03-20 更新2026-03-14 收录
官方服务:

资源简介:

#### Overview This synchronized multimodal neuroimaging dataset for studying brain language processing (SMN4Lang) contains: 1. fMRI and MEG data collected on the same 12 participant while they were listening to 6 hours of naturalistic stories; 2. high-resolution structural (T1, T2), diffusion MRI and resting-state fMRI data for each participant; 3. rich linguistic annotations for the stimuli, including word frequencies, part-of-speech tags, syntactic tree structures, time-aligned characters and words, various types of word and character embeddings. More details about the dataset are described as follows. #### Participants All 12 participants were recruited from universities in Beijing, of which 4 were female, and 8 were male, with an age range 23-30 year. They completed both fMRI and MEG visits (first completed fMRI then MEG experiments which had a gap of 1 month at least), All participants were right-handed adults with Mandarin Chinese as native language who reported having normal hearing and no history of neurological disorders. They were paid and gave written informed consent. The study was conducted under the approval of the Institutional Review Board of Peking University. #### Experimental Procedures Before each scanning, participants first completed a simple information survey form and an informed consent. During both fMRI and MEG scanning, participants were instructed to listen and pay attention to the story stimulus, remain still, answer questions on the screen after each audio was finished. Stimulus presentation was implemented using Psychtoolbox-3. Specifically, at the beginning of each run, there was instruction of "Waiting for the scanning" on the screen followed with 8 seconds blank. Then, the instruction became "This audio is about to start, please listen carefully" which lasted for 2.65 seconds before playing the audio; during audio play, a centrally located fixation cross was presented; finally, two questions about the story were presented each with four answers to choose from during which time was controlled by participants. Auditory story stimuli were delivered via S14 insert earphones for fMRI studies (with headphones or foam padding were placed over the earphones to reduce scanner noise) and Elekta matching insert earphones for MEG studies. The fMRI recording was split into 7 visits with each lasting 1.5 hours in which the T1, T2, resting MRI were collected on the first visit, fMRI with listening tasks was collected from 1 to 6 visits, and the diffusion MRI were collected on the last visit. During MRI scanning including T1, T2, diffusion and resting, participants were instructed to lie relaxed and still in the machine. The MEG recording was split into 6 visits with each lasting 1.5 hours. #### Stimuli Stimuli are 60 story audios with 4 to 7 minutes long, comprising various topics such as education and culture. All audios were downloaded from Renmin Daily Review website read by the same male broadcaster. The corresponding texts were also downloaded from the Renmin Daily Review website in which errors were manually corrected to make sure audio and texts are aligned. #### Annotations Rich annotations of audios and texts are provided in the derivatives/annotations folder, including: 1. Speech to text alignment: The onset and offset time of each character and words in the audio are provided in the "stimuli/time_align" folder. Note that the onset and offset time were added by 10.65 seconds to align with the time of fMRI images because the fMRI scan was started 10.65 seconds before playing the audio. 2. Frequency: Character and word frequencies in the "stimuli/frequency" folder were calculated from the Xinhua news corpus and then log-transformed. 3. Textual embeddings: Text embeddings computed by different pre-trained language models (including Word2Vec, BERT, and GPT2) are provided in the "stimuli/embeddings" folder. Both the character-level and word-level embeddings computed by Word2Vec and BERT model and the word-level embeddings computed by GPT2 model are provided. 4. Syntactic annotations: The POS tag of each word, the constituent tree structure, and the dependency tree structure are provided in the "stimuli/syntactic_annotations" folder. The POS tags were annotated by experts following criterion of Peking Chinese Treebank. The constituent tree structure was manually annotated by linguistic students following PKU Chinese Treebank criterion with the TreeEditor tools and all results were double checked by different experts. The dependency tree structure was transformed from the constituent tree using Stanford CoreNLP tools. #### Preprocessing The MRI data, including the structural, functional, resting and diffusion images, were preprocessed using the “minimal preprocessing pipelines (HCP)” . The MEG data was first preprocessed using the temporal Signal Space Separation (tSSS) method and the bad channels were excluded. And then the independent component analysis (ICA) method was applied to remove the ocular artefacts using the MNE software. #### Usage Notes For the MEG data of sub-08_run-16 and sub-09_run-7, the stimuli-starting triggers were not recorded due to technical problems. The first trigger in these two runs were the stimuli-ending triggers and the starting time can be computed by subtracting the stimuli duration from the time point of the first trigger.

#### 数据集概述 本数据集为用于脑语言加工研究的同步多模态神经影像数据集(SMN4Lang),包含以下内容: 1. 在12名相同被试身上采集的功能磁共振成像(fMRI)与脑磁图(MEG)数据,被试聆听总时长6小时的自然故事; 2. 每名被试的高分辨率结构像(T1、T2)、弥散磁共振成像(diffusion MRI)及静息态功能磁共振成像(resting-state fMRI)数据; 3. 刺激材料附带丰富的语言标注,包括词频、词性标注、句法树结构、时间对齐的字符与词,以及多种类型的词与字符嵌入表示。 有关本数据集的更多细节如下所述。 #### 被试信息 本研究共招募12名被试,均来自北京各高校,其中女性4名,男性8名,年龄区间为23~30岁。所有被试均完成了功能磁共振成像与脑磁图扫描环节,且先完成功能磁共振成像扫描,再进行脑磁图扫描,两次扫描间隔至少1个月。所有被试均为右利手成年人,母语为汉语普通话,自述听力正常,无神经系统疾病史。被试均获得相应报酬,并签署了书面知情同意书。本研究经北京大学伦理审查委员会批准实施。 #### 实验流程 每次扫描前,被试均需填写简易信息调查表并签署知情同意书。功能磁共振成像与脑磁图扫描过程中,要求被试聆听并专注于故事刺激材料,保持静止,并在每段音频播放结束后回答屏幕上呈现的问题。刺激材料的呈现采用Psychtoolbox-3工具实现。具体而言,每个扫描序列开始时,屏幕上将显示"等待扫描"字样,随后呈现8秒的空白画面;随后显示"本次音频即将开始,请仔细聆听"的提示,持续2.65秒后开始播放音频;音频播放期间,屏幕中央呈现注视十字;最后,屏幕将呈现两道与故事内容相关的题目,每题配有四个选项,答题时长由被试自行控制。听觉故事刺激材料通过S14插入式耳机传递:功能磁共振成像扫描时,耳机外需佩戴头戴式耳机或泡沫垫以降低扫描噪音;脑磁图扫描时则采用Elekta配套插入式耳机。 功能磁共振成像扫描共分为7个环节,每个环节持续1.5小时:首个环节采集T1、T2结构像及静息态磁共振成像数据,第1至6个环节采集伴随听觉任务的功能磁共振成像数据,最后一个环节采集弥散磁共振成像数据。在结构像、弥散成像及静息态磁共振扫描过程中,要求被试放松并保持静止躺在扫描仪内。脑磁图扫描共分为6个环节,每个环节持续1.5小时。 #### 刺激材料 本数据集包含60段故事音频,每段时长4至7分钟,涵盖教育、文化等多个主题。所有音频均下载自人民网评论(Renmin Daily Review)网站,由同一名男性播音员录制。对应的文本同样下载自该网站,且已由人工修正其中的错误,以确保音频与文本时间对齐。 #### 标注信息 音频与文本的丰富标注信息存储于derivatives/annotations文件夹中,具体包括: 1. 语音-文本对齐:音频中每个字符与词的起始、结束时间存储于"stimuli/time_align"文件夹中。请注意:由于功能磁共振成像扫描在音频播放前10.65秒即已启动,因此所有标注的起始与结束时间均已增加10.65秒,以与功能磁共振成像图像的时间轴对齐。 2. 词频:"stimuli/frequency"文件夹中的字符与词频数据源自新华社新闻语料库(Xinhua news corpus),且已进行对数变换处理。 3. 文本嵌入表示:"stimuli/embeddings"文件夹中提供了由多种预训练语言模型(包括Word2Vec、BERT及GPT2)计算得到的文本嵌入表示,其中包含Word2Vec与BERT模型生成的字符级、词级嵌入,以及GPT2模型生成的词级嵌入。 4. 句法标注:"stimuli/syntactic_annotations"文件夹中提供了每个词的词性标注、成分句法树结构及依存句法树结构。词性标注遵循北京大学汉语树库(PKU Chinese Treebank)标准,由专业人员完成标注;成分句法树结构由语言学专业学生基于北京大学汉语树库标准,使用TreeEditor工具手动标注,且所有结果均经不同专家进行双重校验;依存句法树结构则通过Stanford CoreNLP工具从成分句法树转换得到。 #### 预处理流程 包括结构像、功能像、静息态成像及弥散成像在内的磁共振成像数据,均采用"最小预处理流程(HCP)"进行预处理。 脑磁图数据首先采用时域信号空间分离(tSSS)方法进行预处理,并剔除坏通道;随后使用MNE软件,通过独立成分分析(ICA)方法去除眼动伪影。 #### 使用注意事项 针对sub-08_run-16与sub-09_run-7的脑磁图数据,由于技术问题,未记录刺激起始触发信号。这两个扫描序列的首个触发信号为刺激结束触发信号,刺激起始时间可通过首个触发信号的时间点减去刺激时长计算得到。

创建时间:
2022-03-20
搜集汇总
数据集介绍
A synchronized multimodal neuroimaging dataset to study brain language processing 数据集图片
背景与挑战
背景概述
该数据集是一个同步多模态神经影像数据集,专门用于研究大脑语言处理。它包含12名汉语母语参与者在聆听6小时自然故事时的fMRI和MEG数据,并提供了高分辨率结构MRI、扩散MRI和静息态fMRI数据。此外,数据集附带了丰富的语言刺激注释,包括词频、句法结构和词嵌入等,支持深入的神经语言学分析。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务