遇见数据集

HUFS-DILAB/MT-wmt14-500k-opus-mt-en-de

收藏
Hugging Face2026-05-13 更新2026-05-31 收录
官方服务:

资源简介:

该数据集基于WMT14英语-德语数据集的训练分割(包含50万句子),源文本为英语句子,参考文本为原始德语参考译文。数据集包含5个翻译候选(h1至h5),每个候选由Helsinki-NLP/opus-mt-en-de模型生成,并带有对数概率分数,按分数降序排序。翻译方向为英语到德语,使用波束搜索方法,参数设置为波束数5和返回序列数5。

This dataset is based on the WMT14 English-German dataset (train split, 500k sentences), with source text in English and reference text as the original German reference. It includes 5 translation candidates (h1 to h5), each generated by the Helsinki-NLP/opus-mt-en-de model and accompanied by log-probability scores, sorted in descending order. The translation direction is English to German, using beam search with parameters num_beams=5 and num_return_sequences=5.

提供机构:
HUFS-DILAB
二维码
社区交流群
二维码
科研交流群
商业服务