HUFS-DILAB/MT-wmt14-500k-opus-mt-en-de
收藏官方服务:
资源简介:
该数据集基于WMT14英语-德语数据集的训练分割(包含50万句子),源文本为英语句子,参考文本为原始德语参考译文。数据集包含5个翻译候选(h1至h5),每个候选由Helsinki-NLP/opus-mt-en-de模型生成,并带有对数概率分数,按分数降序排序。翻译方向为英语到德语,使用波束搜索方法,参数设置为波束数5和返回序列数5。
This dataset is based on the WMT14 English-German dataset (train split, 500k sentences), with source text in English and reference text as the original German reference. It includes 5 translation candidates (h1 to h5), each generated by the Helsinki-NLP/opus-mt-en-de model and accompanied by log-probability scores, sorted in descending order. The translation direction is English to German, using beam search with parameters num_beams=5 and num_return_sequences=5.
提供机构:
HUFS-DILAB


