遇见数据集

EEG Dataset for Second Language Acquisition Arbic and Hindi

收藏
Mendeley Data2026-07-02 收录
官方服务:

资源简介:

This dataset contains electroencephalography (EEG) recordings from 20 healthy adult volunteers — 10 Yemeni native Arabic speakers and 10 Indian native Hindi speakers — during a controlled second language acquisition (L2) task. EEG was recorded using a 40-electrode Virgo EEG system (Allengers), with electrode impedance kept below 15 kΩ, positioned according to the international 10–20 system, and sampled at 256 Hz. Stimuli consisted of 24 single words (12 Arabic + 12 Hindi, semantically matched) presented visually for 10 s each with a 5 s rest interval between sessions. The dataset is released in two complementary forms: (1) RAW — 300 European Data Format (.edf) files, organised per participant and per language. Each participant has its own folder containing an "Arabic" and a "Hindi" subfolder. File-naming convention: S<participant>_<language>_SEE<session>.edf, e.g. S4_A_SEE1.edf (participant 4, Arabic, session 1) or S4_H_SEE4.edf (participant 4, Hindi, session 4). Each participant completed 15 sessions in total, split between the two languages. (2) PROCESSED — a single consolidated CSV file (~1.47 GB) ready for machine-learning pipelines, containing 18 columns: 17 EEG channels (FP1, FP2, F7, F3, F4, F8, A1, T3, CZ, T4, A2, T5, P3, P4, T6, O1, O2) followed by a "label" column indicating the stimulus language (Arabic / Hindi). Participants: 12 male, 8 female; ages 18–46 (mean 29.26 ± 7.72). All participants reported normal hearing, no chronic disease or neurological disorder, and provided written informed consent for public release of de-identified EEG data. Recordings were performed at Medicover Hospital, Aurangabad, Maharashtra, India, between 26 May 2022 and 09 August 2022. Ethics approval: Department of Computer Science and IT, Dr. Babasaheb Ambedkar Marathwada University, Aurangabad, in cooperation with Medicover Hospital (Ref. 91/136/2020). The dataset supports research in neurolinguistics, brain–computer interfaces, EEG signal processing, feature engineering, and machine/deep learning for cross-language neural pattern recognition, with a unique focus on Arabic vs. Hindi as L2. If you use this dataset, please cite BOTH of the following: Aldhaheri, T. A., Kulkarni, S. B., Bhise, P. R., & Tawfik, M. (2025). Utilizing machine and deep learning algorithms to identify learning-related features in electroencephalography data during second language acquisition. Cogent Arts and Humanities, 12(1). https://doi.org/10.1080/23311983.2025.2485696 Aldhaheri, T. A., Kulkarni, S. B., & Al-Zidi, N. M. (2026). Optimizing machine learning models with multi-feature selection for EEG analysis in second language acquisition research. Discover Artificial Intelligence, 6, 99. https://doi.org/10.1007/s44163-025-00801-z

本数据集包含20名健康成年志愿者的脑电图(electroencephalography, EEG)记录:其中10名为以阿拉伯语为母语的也门受试者,10名为以印地语为母语的印度受试者,数据采集于受试者完成受控第二语言习得(second language acquisition, L2)任务期间。 脑电图数据采用40通道Virgo脑电系统(Allengers公司)采集,电极阻抗控制在15 kΩ以下,依照国际10–20系统放置电极,采样率为256 Hz。实验刺激为24个单词语料(12个阿拉伯语单词与12个印地语单词,语义匹配),以视觉方式呈现,每个单词呈现时长为10秒,实验段间设置5秒的休息间隔。 本数据集以两种互补形式发布: (1) 原始数据(RAW):共计300个欧洲数据格式(European Data Format, .edf)文件,按受试者与语言分类组织。每名受试者对应一个专属文件夹,内含“阿拉伯语”与“印地语”两个子文件夹。文件命名规则为:S<受试者编号>_<语言缩写>_SEE<实验段编号>.edf,例如S4_A_SEE1.edf代表受试者4、阿拉伯语任务、第1实验段;S4_H_SEE4.edf代表受试者4、印地语任务、第4实验段。每名受试者共完成15个实验段,分属两种语言任务。 (2) 预处理数据(PROCESSED):单个整合后的CSV文件(约1.47 GB),可直接用于机器学习流程,共包含18列数据:前17列为脑电图通道(FP1、FP2、F7、F3、F4、F8、A1、T3、CZ、T4、A2、T5、P3、P4、T6、O1、O2),最后一列为“标签”列,用于标注刺激语言(阿拉伯语/印地语)。 受试者基本信息:共计12名男性、8名女性,年龄区间为18–46岁(平均年龄29.26 ± 7.72岁)。所有受试者均报告听力正常,无慢性疾病或神经系统病史,并签署书面知情同意书,同意公开其去标识化的脑电数据。 数据采集于2022年5月26日至2022年8月9日期间,在印度马哈拉施特拉邦奥兰加巴德的Medicover医院完成。伦理审批由奥兰加巴德的巴巴萨赫布·阿姆贝德卡尔马哈拉施特拉大学计算机科学与信息技术系与Medicover医院联合批准(审批编号:91/136/2020)。 本数据集可支持神经语言学、脑机接口、脑电信号处理、特征工程,以及跨语言神经模式识别相关的机器学习/深度学习研究,其独特之处在于聚焦阿拉伯语与印地语作为第二语言的研究场景。 若使用本数据集,请同时引用以下两篇文献: Aldhaheri, T. A., Kulkarni, S. B., Bhise, P. R., & Tawfik, M. (2025). Utilizing machine and deep learning algorithms to identify learning-related features in electroencephalography data during second language acquisition. *Cogent Arts and Humanities*, 12(1). https://doi.org/10.1080/23311983.2025.2485696 Aldhaheri, T. A., Kulkarni, S. B., & Al-Zidi, N. M. (2026). Optimizing machine learning models with multi-feature selection for EEG analysis in second language acquisition research. *Discover Artificial Intelligence*, 6, 99. https://doi.org/10.1007/s44163-025-00801-z

创建时间:
2026-05-26
二维码
社区交流群
二维码
科研交流群
商业服务