遇见数据集

One Size Does Not Fit All: Why EU Legislative Translation Demands Domain-Specific Fine-Tuning of LLMs

收藏
Zenodo2026-05-07 更新2026-05-26 收录
官方服务:

资源简介:

This dataset provides two parallel corpora: legislative and non-legislative (generic institutional), covering English as a source language and 23 EU target languages, extracted from the European Parliament's translations stored in the EU interinstitutional translation memory repository Euramis. Each corpus contains 598,000 segment pairs (26,000 per language), split into train (20,000), evaluation (2,000), and test (4,000) sets. The non-legislative corpus is distribution-matched to the legislative one for segment length, and both undergo a four-stage filtering pipeline including exact and near-duplicate removal and semantic similarity-based data leakage detection. Segments are provided in JSONL format with document-level metadata. The dataset was developed within the AI4TRAD project at the European Parliament's Directorate-General for Translation and used in the experiments reported in "One Size Does Not Fit All: Why EU Legislative Translation Demands Domain-Specific Fine-Tuning of LLMs" (EAMT 2026).

提供机构:
Zenodo
创建时间:
2026-05-07
二维码
社区交流群
二维码
科研交流群
商业服务