遇见数据集

DIHUNAMB ! – Breton Periodical, Issue No. 09 (1906) – Parallel Corpus (Breton Vannetais / French)

收藏
Zenodo2026-04-05 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains a sentence-aligned parallel corpus in Breton Vannetais (Gwenedeg) and French, derived from the Breton periodical *DIHUNAMB !* (Issue No. 09, 1906), edited by Loeiz Herrieu and Andréu Mellag. The corpus is part of the IAgwened project, an initiative dedicated to the development of open linguistic resources for the Breton language, with a particular focus on the Vannetais dialect. The dataset has been prepared for use in corpus linguistics, natural language processing (NLP), machine translation, and speech technologies. It follows a structured alignment format (Breton sentence | French sentence) and is encoded in UTF-8. This resource belongs to the HERITAGE corpus category within the IAgwened methodological framework. HERITAGE corpora aim to preserve historical texts with minimal intervention, maintaining original orthography, syntax, and stylistic variation. Three versions of the corpus are provided:- FINAL_DATA: source-aligned corpus preserving the original text- FINAL_ANNOTATED: version enriched with linguistic and contextual annotations- FINAL_NORMALIZED: orthographically normalized version for computational use The dataset is released under the Creative Commons Attribution 4.0 International (CC-BY 4.0) license.

本数据集包含布列塔尼语瓦讷方言(Gwenedeg)与法语的句对齐平行语料库,其源自布列塔尼语期刊《DIHUNAMB !》1906年第09期,由洛埃兹·埃里约(Loeiz Herrieu)与安德烈·梅拉格(Andréu Mellag)编辑。 该语料库属于IAgwened项目的组成部分,该项目致力于开发布列塔尼语的开放语言资源,尤其聚焦于瓦讷方言。 本数据集专为语料库语言学、自然语言处理(Natural Language Processing, NLP)、机器翻译及语音技术研发而制备,采用结构化对齐格式(布列塔尼语句子 | 法语句子),编码为UTF-8。 本资源隶属于IAgwened方法论框架下的HERITAGE语料库类别。HERITAGE语料库旨在以最小干预保留历史文本,维持原始正字法、句法及文体变体。 该语料库提供三个版本: - FINAL_DATA:保留原始文本的源对齐语料库 - FINAL_ANNOTATED:附带语言与上下文标注的增强版本 - FINAL_NORMALIZED:为计算应用进行正字法规范化的版本 本数据集采用知识共享署名4.0国际许可协议(Creative Commons Attribution 4.0 International, CC-BY 4.0)发布。

提供机构:
Zenodo
创建时间:
2026-04-05
二维码
社区交流群
二维码
科研交流群
商业服务