cafeolee/ES-AST_Parallel_Corpus
收藏资源简介:
ES-AST平行语料库是一个西班牙语-阿斯图里亚斯语数据集,旨在支持西班牙资源不足语言(如阿斯图里亚斯语)在自然语言处理任务中的使用,特别是机器翻译。该数据集包含西班牙语和阿斯图里亚斯语的平行句子,可用于训练双向双语机器翻译模型以及多语言机器翻译模型。数据集以两个独立的文本文件(es-ast_corpus.es和es-ast_corpus.ast)和parquet格式(es-ast_corpus.parquet)提供,其中parquet文件包含两列平行文本,每行代表一对平行句子。数据集包含单一训练分割,未应用基于对齐分数的过滤,因此可能包含对齐不良的句子。数据来源包括真实的OPUS平行数据和基于PILAR数据集生成的合成数据,使用Apertium规则翻译器创建。
The ES-AST Parallel Corpus is a Spanish-Asturian dataset created to support the use of under-resourced languages from Spain, such as Asturian, in NLP tasks, specifically Machine Translation. The dataset contains parallel sentences in Spanish and Asturian, and can be used to train Bilingual Machine Translation models between Asturian and Spanish in any direction, as well as Multilingual Machine Translation models. It is provided as two separate txt files (es-ast_corpus.es and es-ast_corpus.ast) and in parquet format (es-ast_corpus.parquet), with the parquet file containing two columns of parallel text where each row represents a pair of parallel sentences. The dataset has a single train split, and no filtering based on alignment score was applied, so it may contain poorly aligned sentences. It aggregates both authentic parallel data from OPUS and synthetic data generated from the Asturian monolingual corpus of the PILAR dataset using the rule-based Apertium translator.



