2024 machine translation evaluation dataset
收藏资源简介:
The dataset was prepared for a LLM to classic NMT evaluation task. The experiment evaluated translation quality between 7 different MT and LLM systems performing the translation task, translating from English to Slovenian language. The original source text was collected in the same day as all the translations were executed, so there was no way for the source text to leak into the observed systems. The observed systems: MT systems: Google Translate: https://translate.google.com/ DeepL Translate~\footnote{DeepL Translate: https://www.deepl.com/en/translator SYSTRAN Translate https://www.systransoft.com/ LLM AI assistants: OpenAI ChatGPT: \url{https://chatgpt.com/ Google Gemini: \url{https://gemini.google.com/app Anthropic Claude: \url{https://claude.ai/ Mistral AI Le Chat: \url{https://chat.mistral.ai/chat The dataset iz available in one table - MS Excel format.



