HalloMTBench
收藏资源简介:
HalloMTBench是一个多语言、人工验证的基准数据集,旨在挑战和诊断现代大型语言模型(LLMs)的翻译幻觉。该数据集包含5,435个高质量的翻译实例,涵盖了11个从英语到其他语言的翻译方向。数据集的创建过程包括使用四种前沿LLMs生成候选翻译,通过集成LLM法官方法进行过滤,并最终由专家进行验证。该数据集可用于评估LLMs在不同语言对上的翻译能力,并揭示模型在翻译幻觉方面的弱点。
HalloMTBench is a multilingual, human-validated benchmark dataset designed to challenge and diagnose translation hallucinations in modern large language models (LLMs). This dataset contains 5,435 high-quality translation instances spanning 11 translation directions from English to other languages. The dataset was developed through a workflow that includes generating candidate translations using four state-of-the-art LLMs, filtering the candidates via an integrated LLM judge-based method, and ultimately validating the results via expert reviews. This benchmark can be used to evaluate the translation capabilities of LLMs across various language pairs, as well as uncover the vulnerabilities of models regarding translation hallucinations.




