surrey-nlp/IndicQE-APE
收藏资源简介:
IndicQE-APE是一个统一的、多任务的质量评估(QE)和自动后编辑(APE)基准数据集,专注于印度语系和其他低资源语言对。该数据集整合了Surrey-NLP QE/APE数据谱系——包括WMT QE直接评估数据、特定领域的印度语QE数据以及人工后编辑数据——形成一个单一的、可配置的资源。每个实例为一行数据,携带每个来源提供的注释(如直接评估分数、后编辑、领域标签),并具有明确的来源信息、确定性内容哈希标识符和每个实例的难度诊断。与随机训练/测试分割不同,该数据集的测试集是首先精心策划的,作为一个困难的诊断工件:它过度代表了操作上困难的实例(如质量有争议、流畅但错误的翻译、错误密集的片段),因此报告的相关性反映了在QE/APE实际重要场景下的性能。数据集包含8种语言对(5种英语到印度语对和3种其他语言到英语对),总实例数为116,332,其中101,552个带有直接评估标签,42,588个带有后编辑。数据集提供多种配置,如qe-da(用于句子级QE)、ape(用于APE)、challenge(诊断测试集)等。数据字段包括源句、机器翻译、后编辑、直接评估分数等。数据集旨在用于评估句子级QE和APE模型,并通过难度分层的挑战集研究鲁棒性。
IndicQE-APE is a consolidated, multi-task Quality Estimation (QE) and Automatic Post-Editing (APE) benchmark for Indic and other low-resource language pairs, with a test-first, difficulty-stratified challenge split. It unifies the Surrey-NLP QE/APE data lineage — including WMT QE Direct-Assessment data, domain-specific Indic QE, and human post-edits — into a single, configurable resource. Each instance is one row carrying whatever annotation each source provides (such as DA scores, post-edits, domain labels), with explicit provenance, a deterministic content-hash identifier, and per-instance difficulty diagnostics. Unlike a random train/test split, the test set is curated first as a hard, diagnostic artifact: it over-represents the operationally difficult instances (e.g., contested quality, fluent-but-wrong translations, error-dense segments) so reported correlations reflect performance where QE/APE actually matters. The dataset includes 8 language pairs (5 English-to-Indic curated pairs and 3 other language-to-English passthrough pairs), with a total of 116,332 instances, of which 101,552 are QE-labelled with Direct Assessment and 42,588 have post-edits. It offers multiple configurations, such as qe-da (for sentence-level QE), ape (for APE), and challenge (the diagnostic test set). Data fields include source sentence, machine translation, post-edit, DA scores, etc. The dataset is intended for benchmarking sentence-level QE and APE models and studying robustness via the difficulty-stratified challenge set.




