AnglistikVoices: L2 English speech dataset
收藏资源简介:
AnglistikVoices: an L2 English speech dataset This repository contains an L2 (second language) English speech corpus consisting of 74 minutes of recorded audio from 15 non-native English speaking participants. The dataset was created as part of a university course, with all participants being students who are also the authors of this dataset. Dataset Specifications Total participants: 15 non-native English speakers Total audio duration: 74 minutes Recordings per participant: 60 audio samples each Sentence alignment: Available for 8 out of 15 participant Recording equipment: Audio-Technica ATM75 microphone Stimuli: All sentences are from the Artie Bias Corpus (https://github.com/artie-inc/artie-bias-corpus) Recording environment: Recording booth The dataset contains individual recordings of non-native English speakers organized by participant ID. For 8 participants, sentence-level alignments are provided. All recordings were captured in a controlled acoustic environment using Audio-Technica ATM75 microphone to ensure high audio quality. The recordings consist of spoken English utterances from each participant. Detailed linguistic profiles for each participant are available in the metadata.xlsx file, which is indexed by participant ID and contains information on native language, proficiency level, language learning history, and other relevant linguistic background data. The audio files are organized by participant ID, matching the identifiers used in the metadata file for easy cross-referencing between the audio recordings and participant linguistic profiles. Authors and Contributors This dataset was created by the student participants themselves as part of their coursework. Course Instructor: Akhilesh Kakolu Ramarao Teaching Assistant: Anna Sophia Stein If you have any questions, you can contact: kakolura@hhu.de If you use this dataset in your research, please cite: @dataset{kakolu_ramarao_2024_anglistikvoices, author = {Kakolu Ramarao, A. and Stein, A. S. and Tahiri, A. and Rodrigues, D. C. and Antonia Weismann, C. and Schäfer, O. S. and Kaczor, J. and Tran, N. H. and Elena Telaar, C. and Bauer, L. and Jütten, M. and Mafuta, C. and Agelopoulou, V. V. and Grabowski, Q. A. G.}, title = {AnglistikVoices: L2 English speech dataset}, publisher = {Zenodo}, version = {v1.0.0}, year = {2024}, month = jun, doi = {10.5281/zenodo.12525952}, url = {https://doi.org/10.5281/zenodo.12525952}, note = {LabPhon 19, Hanyang Institute for Phonetics and Cognitive Sciences of Language (HIPCS), Hanyang University in Seoul, Korea} }
AnglistikVoices:二语英语语音数据集 本仓库收录一套第二语言(L2)英语语音语料库,包含来自15名非英语母语使用者的共计74分钟录音音频。本数据集依托一门大学课程项目开发,所有参与者均为课程学生,同时也是本数据集的作者。 数据集规格参数 总参与人数:15名非英语母语使用者 总音频时长:74分钟 单参与者录音量:每名参与者提供60条音频样本 句子级对齐标注覆盖情况:15名参与者中,8名提供了句子级对齐标注 录音设备:铁三角ATM75麦克风(Audio-Technica ATM75) 刺激语句源:所有语句均取自Artie偏差语料库(Artie Bias Corpus,https://github.com/artie-inc/artie-bias-corpus) 录音环境:隔音录音棚 本数据集按参与者ID对非英语母语使用者的独立录音进行组织。其中8名参与者配有句子级对齐标注。所有录音均在可控声学环境中使用铁三角ATM75麦克风采集,以保障优异的音频质量。 录音内容为每名参与者的英语口语话语。每名参与者的详细语言特征档案可在metadata.xlsx文件中获取,该文件以参与者ID为索引,涵盖母语信息、语言熟练度等级、语言学习经历及其他相关语言背景数据。 音频文件同样按参与者ID进行组织,与元数据文件中使用的标识符一一对应,便于音频录音与参与者语言特征档案之间的交叉检索与关联分析。 作者与贡献者 本数据集由参与课程的学生作为课程项目的一部分独立创建。 课程主讲教师:Akhilesh Kakolu Ramarao 教学助理:Anna Sophia Stein 如有任何疑问,可联系:kakolura@hhu.de 若在研究工作中使用本数据集,请引用如下: @dataset{kakolu_ramarao_2024_anglistikvoices, author = {Kakolu Ramarao, A. and Stein, A. S. and Tahiri, A. and Rodrigues, D. C. and Antonia Weismann, C. and Schäfer, O. S. and Kaczor, J. and Tran, N. H. and Elena Telaar, C. and Bauer, L. and Jütten, M. and Mafuta, C. and Agelopoulou, V. V. and Grabowski, Q. A. G.}, title = {AnglistikVoices: L2 English speech dataset}, publisher = {Zenodo}, version = {v1.0.0}, year = {2024}, month = jun, doi = {10.5281/zenodo.12525952}, url = {https://doi.org/10.5281/zenodo.12525952}, note = {LabPhon 19, Hanyang Institute for Phonetics and Cognitive Sciences of Language (HIPCS), Hanyang University in Seoul, Korea} }



