Speak & Improve (S&I) Corpus 2025
收藏资源简介:
Speak & Improve (S&I) Corpus 2025是由剑桥大学ALTA研究所/MIL实验室创建的一个用于研究口语语言评估和反馈的数据集。该数据集包含约340小时的第二语言学习者英语口语数据,涵盖了从A2到C1的欧洲语言共同参考框架(CEFR)水平。数据集内容包括详细的转录、不流畅标签、语法错误修正和熟练度评分,旨在支持口语语言评估和反馈的研究。数据集的创建过程包括三个阶段:评分、转录标注和错误标注,确保了数据的高质量和实用性。该数据集主要应用于语言学习领域,旨在解决自动口语评估和反馈中的技术挑战。
Speak & Improve (S&I) Corpus 2025 is a dataset developed by the ALTA Institute / MIL Lab at the University of Cambridge for research on spoken language assessment and feedback. It contains approximately 340 hours of spoken English data from second language learners, spanning CEFR proficiency levels from A2 to C1. The dataset includes detailed transcripts, disfluency labels, grammatical error corrections, and proficiency ratings, and was constructed via three stages: scoring, transcription annotation, and error annotation to ensure its high quality and practical applicability. Primarily utilized in the field of language learning, this dataset aims to address technical challenges in automatic spoken language assessment and feedback.




