LearnerVoice
收藏资源简介:
LearnerVoice数据集由韩国科学技术院计算学院创建,包含50.04小时的非母语英语学习者自发语音数据,共计229,671个token。数据集通过在线学习平台Ringle收集,该平台提供一对一视频辅导课程。数据集的创建过程中,特别关注了非母语学习者语音中的不规则表达和断句特征,并由专业标注人员进行详细转录。该数据集主要用于改进自动语音识别系统,特别是在处理非母语英语学习者的自发语音时,提高识别准确性和流畅性。
The LearnerVoice dataset was created by the School of Computing at Korea Advanced Institute of Science and Technology (KAIST). It contains 50.04 hours of spontaneous speech data from non-native English learners, totaling 229,671 tokens. The dataset was collected via the online learning platform Ringle, which offers one-on-one video tutoring sessions. Special attention was paid to the irregular expressions and pause/sentence-breaking patterns in the speech of non-native learners during the dataset construction, and detailed transcriptions were performed by professional annotators. This dataset is primarily used to improve automatic speech recognition (ASR) systems, particularly to enhance recognition accuracy and fluency when handling spontaneous speech from non-native English learners.
数据集概述
数据集内容
- 数据收集: 包含从学生与母语者实时互动中收集的广泛数据集,包括详细的语言注释、熟练度评估和用户反馈。
- AI语言模型: 涉及我们专有的AI系统的内部工作原理,包括算法性能指标、匿名用户交互数据和模型训练集。
访问数据
- 学术成员: 学术机构或研究组织的成员可以注册账户,获得对数据集和AI工具的完全访问权限。
- 非成员: 访客可以申请有限访问权限,需提交研究需求和数据使用意图的详细说明。
数据许可
- 特定数据集和工具可供许可,请联系我们的许可部门了解条款和条件。
出版物
- Ringle的见解: 我们的出版物专注于语言学习的深入分析和研究成果。




