FCE
收藏资源简介:
displayName: FCE (First Certificate in English) labelTypes: - English Corpus - Classification license: - FCE Custom mediaTypes: - Text paperUrl: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.478.2066&rep=rep1&type=pdf publishDate: "2011" publishUrl: https://ilexir.co.uk/datasets/index.html publisher: - University of Cambridge tags: - Test taskTypes: - Grammatical Error Detection/Correction --- # 数据集介绍 ## 简介 CLC FCE 数据集是一组 1,244 份试卷,由 2000 年和 2001 年参加剑桥 ESOL 第一英语证书 (FCE) 考试的考生编写。 这些脚本是从剑桥学习者语料库 (CLC) 中提取的,该语料库是剑桥大学出版社和剑桥评估公司合作开发的。 对于每个考试脚本,CLC FCE 数据集包括考生编写的原始文本(转录和匿名,但未经修改)以及分数、错误注释和基本人口统计细节,包括考生的第一语言和年龄范围。 ## 引文 ``` @inproceedings{yannakoudakis2011new, title={A new dataset and method for automatically grading ESOL texts}, author={Yannakoudakis, Helen and Briscoe, Ted and Medlock, Ben}, booktitle={Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies}, pages={180--189}, year={2011} } ``` ## Download dataset :modelscope-code[]{type="git"}
displayName: FCE(英语第一证书考试,First Certificate in English) labelTypes: - 英语语料库(English Corpus) - 分类任务(Classification) license: - FCE定制许可(FCE Custom) mediaTypes: - 文本(Text) paperUrl: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.478.2066&rep=rep1&type=pdf publishDate: "2011" publishUrl: https://ilexir.co.uk/datasets/index.html publisher: - 剑桥大学(University of Cambridge) tags: - 考试(Test) taskTypes: - 语法错误检测与纠正(Grammatical Error Detection/Correction) --- # 数据集简介 ## 数据集概况 CLC FCE数据集包含1244份答卷,均由2000年与2001年参加剑桥英语为非母语者(English for Speakers of Other Languages,ESOL)第一英语证书(First Certificate in English,FCE)考试的考生作答完成。 该数据集的文本均源自剑桥学习者语料库(Cambridge Learner Corpus,CLC),该语料库由剑桥大学出版社与剑桥评估公司联合开发。 针对每一份考试答卷,CLC FCE数据集均提供考生作答的原始文本(已完成转录与匿名化处理,但未做任何修改),同时附带考试得分、错误标注以及考生的基础人口统计学信息,其中包含考生的母语与年龄区间。 ## 引用文献 @inproceedings{yannakoudakis2011new, title={面向英语非母语文本自动评分的新型数据集与方法}, author={Yannakoudakis, Helen and Briscoe, Ted and Medlock, Ben}, booktitle={第49届国际计算语言学协会年会:人类语言技术分会场论文集}, pages={180--189}, year={2011} } ## 数据集下载 :modelscope-code[]{type="git"}

- FCE数据集首次发表,作为剑桥大学英语写作评估的一部分,旨在提供一个标准化的英语写作错误标注数据集。
- FCE数据集首次应用于自然语言处理领域,特别是在错误检测和纠正任务中,为研究者提供了一个重要的基准数据集。
- FCE数据集的扩展版本发布,增加了更多的写作样本和详细的错误分类,进一步丰富了数据集的内容和应用范围。
- FCE数据集被广泛应用于机器学习和人工智能领域,特别是在自动作文评分和写作辅助系统中,成为该领域的重要资源。
- 1The FCE corpus: A resource for error detection researchUniversity of Cambridge · 2008年
- 2Automatic Error Detection in Learner Writing: A Large-Scale Multi-Class Classification TaskUniversity of Cambridge · 2019年
- 3Improving Grammatical Error Detection in Essays Using Deep LearningUniversity of Cambridge · 2020年
- 4A Comparative Study of Grammatical Error Detection Systems on the FCE CorpusUniversity of Cambridge · 2018年
- 5Exploring the Use of BERT for Grammatical Error Detection in the FCE CorpusUniversity of Cambridge · 2021年



