遇见数据集

LumiOpen/lingsoft-pemt-evals

收藏
Hugging Face2026-07-07 更新2026-07-22 收录
官方服务:

资源简介:

--- language: - fi - en license: cc-by-4.0 task_categories: - translation - multiple-choice pretty_name: Lingsoft PEMT translation preference evals dataset_info: - config_name: FI_pairs features: - name: id dtype: string - name: task_id dtype: string - name: language_pair dtype: string - name: date dtype: string - name: segment dtype: string - name: segment_no dtype: string - name: source_text dtype: string - name: raw_mt dtype: string - name: target_pemt dtype: string - name: topics dtype: string splits: - name: test_shared num_bytes: 360233 num_examples: 527 - name: test_broad num_bytes: 2709525 num_examples: 4072 download_size: 1195504 dataset_size: 3069758 - config_name: FI_preference_5shot features: - name: id dtype: string - name: pair_id dtype: string - name: language dtype: string - name: split dtype: string - name: order dtype: string - name: shots dtype: int64 - name: prompt dtype: string - name: choices sequence: string - name: answer dtype: string - name: completion dtype: string - name: metadata struct: - name: option_a_source dtype: string - name: option_b_source dtype: string - name: segment dtype: string - name: segment_no dtype: string - name: task_id dtype: string splits: - name: test_shared num_bytes: 3015838 num_examples: 1054 - name: test_broad num_bytes: 23253850 num_examples: 8144 download_size: 2925824 dataset_size: 26269688 configs: - config_name: FI_pairs data_files: - split: test_shared path: FI_pairs/test_shared-* - split: test_broad path: FI_pairs/test_broad-* - config_name: FI_preference_5shot data_files: - split: test_shared path: FI_preference_5shot/test_shared-* - split: test_broad path: FI_preference_5shot/test_broad-* --- # Lingsoft PEMT translation preference evals Evaluation sets for translation quality and fluency, built from the [Lingsoft-EU-Summaries-PEMT](https://github.com/LumiOpen/Lingsoft-EU-Summaries-PEMT) corpus: professional post-edited machine translations of "Summaries of EU Legislation" (2016–2025). Each item pairs a raw machine translation (`Raw_MT`) with its professional human post-edit (`Target_PEMT`) for the same English source sentence. Only pairs where post-editing changed the text are included. Intended for use with the [LumiOpen lm-evaluation-harness](https://github.com/LumiOpen/lm-evaluation-harness) `lingsoft_pemt` tasks. ## Configs ### `FI_preference_5shot` Pre-rendered blind A/B preference prompts for Finnish (task `lingsoft_pemt_fi_mcf_*`): the prompt shows the English source and both translations labeled only "Finnish translation A/B", with 5 few-shot examples (sampled from the FI train split, balanced answer labels) baked into the prompt. Every pair appears twice with the A/B order swapped (`order` = `raw_first` / `target_first`), so random chance is exactly 50%. Score the single-token continuations `" A"` / `" B"` by loglikelihood; the correct answer is always the post-edited side. Generated by `create_fi_base_model_eval.py` in the source repo (5 shots, seed 42). ### `FI_pairs` The underlying changed rows (task `lingsoft_pemt_fi_cf_*`): `source_text`, `raw_mt`, `target_pemt` plus corpus metadata. `id` matches `pair_id` in `FI_preference_5shot`. Use for completion-form scoring — compare loglikelihoods of the two translations directly as continuations of the source. ## Splits - `test_shared`: text units from 6 segments shared across all 23 corpus languages (cross-lingually comparable if more languages are added). - `test_broad`: a broad sample across all segments. | config | test_shared | test_broad | |---|---|---| | FI_preference_5shot (records) | 1,054 | 8,144 | | FI_pairs (pairs) | 527 | 4,072 | ## Provenance and license Derived from professional translation and post-editing projects by Lingsoft under "Summaries of EU Legislation" (2016–2025). Original editorial content © European Union, re-used via the Publications Office of the European Union under CC-BY 4.0. Proprietary triple-stage alignment and raw MT layers provided by Lingsoft.

Lingsoft PEMT translation preference evals dataset is designed for evaluating translation quality and fluency. It is built from the Lingsoft-EU-Summaries-PEMT corpus, consisting of professional post-edited machine translation pairs where each item pairs a raw machine translation (Raw_MT) with its human post-edited version (Target_PEMT) for the same English source sentence, including only pairs where post-editing changed the text. The dataset is intended for use with the LumiOpen lm-evaluation-harness lingsoft_pemt tasks, supporting Finnish and English. It includes two configurations: FI_pairs provides the underlying translation pairs with metadata for completion-form scoring, and FI_preference_5shot offers pre-rendered blind A/B preference prompts with 5 few-shot examples for preference evaluation. Splits include test_shared (text units shared across 23 languages) and test_broad (a broad sample), with examples ranging from 527 to 8,144. Derived from professional translation and post-editing by Lingsoft under Summaries of EU Legislation (2016–2025), it is licensed under CC-BY 4.0.

提供机构:
LumiOpen
搜集汇总
数据集介绍
LumiOpen/lingsoft-pemt-evals 数据集图片
构建方式
本数据集源自Lingsoft公司针对“欧盟立法摘要”(2016–2025年)所开展的专业译后编辑项目,基于其公开的Lingsoft-EU-Summaries-PEMT语料库构建。构建流程首先将每一条英文源句与其对应的原始机器翻译输出及经过人工译后编辑的芬兰语译文进行配对,并仅保留那些编辑确实改变了文本内容的样本对。此外,对于偏好评估配置,系统额外生成了盲测A/B提示,每个译文对以随机顺序出现两次,并嵌入了五个平衡答案标签的少样本示例,从而消除位置偏差,确保评估的公正性。
使用方法
使用本数据集时,研究者需借助LumiOpen的lm-evaluation-harness工具链中的lingsoft_pemt任务进行调用。对于FI_preference_5shot配置,应通过比较模型在选项“A”与“B”上的对数似然得分来判定其偏好,得分较高的选项即代表模型更倾向于经过人工译后编辑的译文。而对于FI_pairs配置,则可直接以源文本为上下文,分别计算原始机器译文与人工编辑译文的续写对数似然,从而在更精细的层面评估翻译质量。所有数据以配置文件形式组织,便于按需加载指定子集。
背景与挑战
背景概述
该数据集由Lingsoft机构于2025年之前创建,依托欧盟立法摘要(2016–2025年)的专业后编辑机器翻译语料库,聚焦芬兰语与英语的翻译质量评估。核心研究问题在于通过成对比较原始机器翻译与人工后编辑译文,系统评估机器翻译的流畅性与忠实度。数据集包含两个配置:FI_pairs提供可直接用于似然度评分的平行句对,FI_preference_5shot则构建了盲测偏好提示,以消除顺序偏差。其双测试集设计(test_shared为跨语言可比样本,test_broad为广泛采样)为多语言翻译评估提供了基准框架,对低资源语言翻译质量研究具有推进意义。
当前挑战
领域挑战在于,机器翻译质量评估长期依赖BLEU等自动指标,难以捕捉语义忠实度与语用流畅性;该数据集通过人类后编辑对比,直接量化了翻译提升空间。构建挑战包括三重阶段的语料对齐复杂性——需统一欧盟多语种法律文本的句段边界,并仅保留后编辑导致文本变异的子集;此外,为避免标注偏差,每对译文以AB顺序轮换呈现,确保50%随机基线,这对提示模板设计与样本平衡提出了精细要求。
常用场景
经典使用场景
在机器翻译(MT)质量评估领域,lingsoft-pemt-evals数据集为评估翻译系统的译后编辑效果提供了宝贵的基准资源。该数据集源自欧盟立法摘要的专业译后编辑流程,通过配对原始机器翻译(Raw_MT)与人工精校后的目标译文(Target_PEMT),构建了精准的偏好测试样本。研究者可借助这些数据开展盲测A/B比较,衡量机器翻译产出在专业编辑后的改善程度,从而为翻译系统的优化提供量化依据。
解决学术问题
该数据集有效解决了低资源语言机器翻译质量评估中的两大关键难题:一是缺乏标准化的、由专业译员后编辑构建的平行偏好数据;二是难以区分模型流畅性与忠实性的改进来源。通过提供芬兰语-英语的高质量译后编辑配对样本,该数据集使研究者能够基于对数似然评分或上下文续接概率,客观比较不同翻译候选与后编辑版本的差异,从而推动统计显著性检验和生成质量归因分析。
实际应用
在实际应用层面,该数据集直接服务于多语言翻译产品的质量监控与迭代。翻译服务提供商可利用其中的偏好评估数据训练翻译记忆库或自动质量评估模型,精准识别机器翻译输出的常见缺陷模式。对于部署了机器翻译系统的企业而言,该数据集有助于构建基于译后编辑成本的量化指标,指导翻译流程中人工干预点的设置与资源分配,最终提升大规模本地化项目的效率与一致性。
数据集最近研究
最新研究方向
该数据集聚焦于机器翻译质量评估与人类偏好对齐的前沿探索,特别是在低资源语言对(如芬兰语-英语)的场景中。通过引入专业后期编辑的翻译对与盲测A/B偏好提示,为评估大语言模型的翻译流畅性与忠实度提供了高标准的基准。当前研究热点包括利用此类细粒度偏好数据优化模型的奖励建模与强化学习微调,以提升输出的语言学自然度与任务符合性。该数据集的跨语言共享子集设计亦为多语言翻译系统的公平性比较与泛化能力研究奠定了基础,为欧洲多语言数字生态中的翻译技术伦理与质量控制提供了关键资源。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务