RockArt-PEARMUT v1: Human evaluation dataset for glossary-augmented LLM and NMT translation of Spanish rock art text
收藏资源简介:
This dataset contains the PEARMUT campaign configurations, raw human annotation export, terminology occurrence files, and derived result summaries used in a human evaluation of Spanish-to-English machine translation for a terminology-dense rock art text. The study compares three MT configurations: DeepL as a commercial NMT baseline, Gemini-Simple as an LLM baseline with a basic translation prompt, and Gemini-RAG as a glossary-augmented LLM configuration using lightweight retrieval of relevant Spanish-English terminology pairs. The dataset supports two complementary evaluation tasks. Task 1 is a multi-way DA-style evaluation in which 91 Spanish source segments are shown with three anonymised English outputs and scored from 0 to 100 for overall translation quality. Task 2 is a terminology-only MQM-style audit in which 194 auto-detected glossary term occurrences are checked across the same three systems using a restricted terminology taxonomy: wrong term, missing term, and inconsistent term. The paper reports 273 segment-system DA ratings and 582 terminology checks. The derived summaries show that Gemini-RAG preserves overall DA-style quality while improving exact-match terminology accuracy. Mean DA scores were 85.27 for Gemini-RAG, 85.24 for Gemini-Simple, and 80.27 for DeepL. Exact-match terminology accuracy was 81.44% for Gemini-RAG, 69.07% for Gemini-Simple, and 64.43% for DeepL. Read the README.md for the instructions.



