ALIA_syntethic_MT_V2
收藏资源简介:
ALIA Synthetic MT V2是一个多语言平行语料库,专门为巴斯克语(Euskera)的机器翻译任务设计。数据集包含三个部分:1) Berria新闻语料,来自Berria新闻网站的337,650篇文档(2023-2025年),提供巴斯克语原文及使用Qwen3-32B和LatxaQ模型生成的英语和西班牙语合成翻译,并包含Sonar相似度评分;2) 巴斯克官方公报语料,包含27,130个法律和法规文本示例,源自巴斯克官方公报,附带丰富元数据和西班牙语翻译;3) 议会演讲语料,包含20,385个议会演讲文本,涵盖元数据信息及西班牙语翻译。数据集使用先进的大语言模型生成高质量合成翻译,其中LatxaQ模型针对巴斯克语言和文化进行了优化。适用于机器翻译训练与评估、合成数据实验及多语言NLP研究,但翻译可能包含模型伪影。采用CC-BY-SA-4.0许可证发布。
ALIA Synthetic MT V2 is a multilingual parallel corpus specifically constructed for machine translation tasks involving Basque (Euskera). The dataset consists of three main components: 1) Berria news corpus, comprising 337,650 documents from the Berria news website (2023-2025), with Basque originals and synthetic translations in English and Spanish generated using Qwen3-32B and LatxaQ large language models, along with Sonar similarity scores for quality assessment; 2) Basque official gazette corpus, containing 27,130 legal and regulatory text examples from the Basque official gazette, with extensive metadata and Spanish translations; 3) Parliament speech corpus, with 20,385 parliamentary speech texts, including metadata and Spanish translations. The dataset leverages advanced LLMs to produce high-quality synthetic translations, with LatxaQ optimized for Basque linguistic and cultural features. It is suitable for machine translation model training and evaluation, synthetic data experiments, and multilingual NLP research, though translations may contain model-specific artifacts. Released under the CC-BY-SA-4.0 license.
数据集概述
- 名称: ALIA Synthetic MT V2 (ALIA_syntethic_MT_V2)
- 许可证: CC-BY-SA 4.0
- 语言: 巴斯克语 (eu)、英语 (en)、西班牙语 (es)
- 任务类别: 翻译 (translation)
- 数据集规模: 100K < n < 1M
数据集组成
该数据集包含三个子集(config):
-
berria_eu_en_es (新闻)
- 来源: Berria 新闻文章 (2023-2025年)
- 翻译方式: 合成翻译(由Qwen3-32B和LatxaQ模型生成)
- 样本数: 337,650
- 数据字段:
id: 文档唯一标识符doc_eu: 巴斯克语原文doc_en: 合成英语翻译doc_es: 合成西班牙语翻译sonar_en_score: 巴斯克语与英语的Sonar相似度分数 (float64)sonar_es_score: 巴斯克语与西班牙语的Sonar相似度分数 (float64)
-
bopv_eu_en_es (法律/官方公报)
- 来源: BOPV (巴斯克官方公报)
- 样本数: 27,130
- 数据字段 (包含元数据与文本):
titulo,normative_range,estado,ambito,numBulletin,numOrder,numDisposal,fechaPublicacion,fechaDisposicion,organismo,departamento,seccion,temas,urleu_unique_id,es_unique_ideu_text: 巴斯克语文本es_text: 西班牙语文本eu_word_count: 巴斯克语词数 (int64)translation: 翻译字符串
-
parliament_eu_en_es (议会记录)
- 来源: 议会会议记录
- 样本数: 20,385
- 数据字段:
legislatura,fecha,speaker,party,topic,language,urleu_unique_id,es_unique_ideu_text: 巴斯克语文本es_text: 西班牙语文本eu_word_count: 巴斯克语词数 (int64)translation: 翻译字符串
数据集统计
| 子集 (config) | 训练集样本数 | 下载大小 | 数据集大小 (字节) |
|---|---|---|---|
| berria_eu_en_es | 337,650 | 315.5 MB | 520.5 MB |
| bopv_eu_en_es | 27,130 | 69.8 MB | 181.2 MB |
| parliament_eu_en_es | 20,385 | 64.5 MB | 123.5 MB |
生成模型
- Qwen3-32B: 多语言大语言模型。
- LatxaQ-VL: 基于Qwen3-32B的领域适应版本,保留了巴斯克语能力,目前为非公开研究模型,计划未来开放。
数据用途
- 机器翻译模型的训练与评估
- 合成数据实验
- 多语言自然语言处理研究
引用与资金
- 引用: 使用该数据集时,请引用ALIA项目及所用翻译模型。
- 资金: 由“Ministerio para la Transformación Digital y de la Función Pública”资助,属于“Desarrollo de Modelos ALIA”项目的一部分,由欧盟NextGenerationEU提供资金支持。




