One Size Does Not Fit All: Why EU Legislative Translation Demands Domain-Specific Fine-Tuning of LLMs
收藏资源简介:
This dataset provides two parallel corpora: legislative and non-legislative (generic institutional), covering English as a source language and 23 EU target languages, extracted from the European Parliament's translations stored in the EU interinstitutional translation memory repository Euramis. Each corpus contains 598,000 segment pairs (26,000 per language), split into train (20,000), evaluation (2,000), and test (4,000) sets. The non-legislative corpus is distribution-matched to the legislative one for segment length, and both undergo a four-stage filtering pipeline including exact and near-duplicate removal and semantic similarity-based data leakage detection. Segments are provided in JSONL format with document-level metadata. The dataset was developed within the AI4TRAD project at the European Parliament's Directorate-General for Translation and used in the experiments reported in "One Size Does Not Fit All: Why EU Legislative Translation Demands Domain-Specific Fine-Tuning of LLMs" (EAMT 2026).



