UniversalCEFR/ComplexityMT
收藏资源简介:
ComplexityMT是一个基准数据集,用于研究文本复杂度(通过欧洲共同语言参考框架(CEFR)操作化)与机器翻译之间的交互作用。该数据集包含来自五个机器翻译系统的输出,这些输出翻译了涵盖所有六个CEFR级别(A1至C2)的文本,涉及六种语言,同时提供了无参考翻译质量分数(COMET-22、GEMBA-DA)以及前后向翻译输出的CEFR分类器读数。数据集支持两个互补任务:任务1(鲁棒性)研究源文本复杂度是否影响翻译质量;任务2(保持性)研究机器翻译是否保持源文本的CEFR水平。数据集覆盖的语言包括阿拉伯语、荷兰语、英语、法语、印地语和俄语。MT系统包括GPT-5.4、Google Translate、TranslateGemma 4B、TranslateGemma 12B和Tower-Instruct 7B。数据集结构包括任务1和任务2的配置,每个配置有句子和文档分割,数据字段包括源语言、目标语言、CEFR级别、翻译文本、质量分数等。创建数据集的目的是评估机器翻译系统在保持源文本教学复杂度方面的表现,源数据来自UniversalCEFR集合,注释包括人类标注的黄金CEFR标签、自动翻译质量分数和CEFR分类器读数。
ComplexityMT is a benchmark dataset for investigating the interplay between text complexity (operationalized via the Common European Framework of Reference for Languages, CEFR) and machine translation. This dataset contains outputs from five machine translation systems, which translated texts spanning all six CEFR proficiency levels (A1 to C2) across six languages. It also provides reference-free translation quality metrics (COMET-22, GEMBA-DA) and CEFR classifier scores for forward and backward translation outputs. The dataset supports two complementary tasks: Task 1 (Robustness) investigates whether source text complexity impacts translation quality; Task 2 (Retention) examines whether machine translation preserves the CEFR proficiency level of the source text. The languages covered in the dataset include Arabic, Dutch, English, French, Hindi, and Russian. The MT systems included are GPT-5.4, Google Translate, TranslateGemma 4B, TranslateGemma 12B, and Tower-Instruct 7B. The dataset structure includes configurations for Task 1 and Task 2, each with sentence-level and document-level splits. The data fields include source language, target language, CEFR proficiency level, translated text, quality metrics, and more. The dataset was created to evaluate machine translation systems' performance in preserving the pedagogical complexity of source texts as defined by their CEFR levels. The source data originates from the UniversalCEFR collection, and the annotations include human-annotated gold-standard CEFR labels, automatic translation quality metrics, and CEFR classifier scores.




