NewsPaLM MBR and QE Dataset
收藏资源简介:
NewsPaLM MBR和QE数据集是由谷歌研究团队开发的,基于PaLM-2 Bison大型语言模型生成的英德和德英平行数据集。该数据集包含句子级和多句子级的示例,通过最小贝叶斯风险(MBR)解码和质量估计(QE)重排序生成。数据集的创建过程包括源侧数据收集、构建“blob”、基于聚类的文本选择以及MBR解码和QE重排序。该数据集主要用于神经机器翻译(NMT)模型的预训练和微调,旨在提高NMT模型的性能,特别是在处理长序列和多句子数据时。
The NewsPaLM MBR and QE Dataset was developed by the Google Research team. It is an English-German and German-English parallel dataset generated using the PaLM-2 Bison Large Language Model. This dataset encompasses sentence-level and multi-sentence-level examples, which are produced through Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) re-ranking. The development pipeline of this dataset includes source-side data collection, constructing "blobs", clustering-based text selection, as well as MBR decoding and QE re-ranking. This dataset is primarily used for pre-training and fine-tuning of Neural Machine Translation (NMT) models, with the goal of improving the performance of NMT models, particularly when dealing with long sequences and multi-sentence data.

- 1Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data谷歌 · 2024年



