HiTZ/ALIA_scientic_articles
收藏资源简介:
该数据集包含一个高质量、真实(非合成)的平行语料库,专注于科学和学术领域,涵盖巴斯克语(eu)、西班牙语(es)和英语(en)。与合成生成的数据集不同,该语料库基于真实、人工创作的三语摘要构建,这些摘要提取自巴斯克大学(UPV/EHU)的机构数字存储库ADDI(Archivo Digital de Docencia e Investigación)。它专门设计用于支持机器翻译系统和大型语言模型(LLM)在专业科学文本上的训练、微调和评估。数据来源包括ADDI存储库中的多个学术集合,文档类型涵盖博士论文、硕士论文、学士论文以及学术期刊和文章。每个平行记录都附有丰富的元数据(如出版日期、提交日期、URI、语言、许可证、标题和部门),以方便针对性过滤。数据集包含3,294个完美对齐的三语文档(即3,294组巴斯克语、西班牙语和英语元组)。巴斯克语的平均单词数低于西班牙语和英语,这反映了巴斯克语的黏着语形态特征。
This dataset contains a high-quality, real (non-synthetic) parallel corpus in the scientific and academic domain, covering Basque (eu), Spanish (es), and English (en). Unlike synthetically generated datasets, this corpus is built from genuine, human-authored trilingual abstracts extracted from ADDI (Archivo Digital de Docencia e Investigación), the institutional digital repository of the University of the Basque Country (UPV/EHU). It is specifically designed to support the training, fine-tuning, and evaluation of Machine Translation systems and Large Language Models (LLMs) for specialized scientific text. The raw documents were scraped and curated from multiple academic collections within the ADDI repository, including Doctoral Theses, Masters Theses, Bachelors Theses, and Academic Journals and Journal Articles. Each parallel record is accompanied by rich metadata (e.g., Publication Date, Submission Date, URI, Language, License, Title, Department) to facilitate targeted filtering. The dataset consists of 3,294 perfectly aligned trilingual documents (yielding 3,294 tuples of Basque, Spanish, and English). The lower average word count in Basque compared to Spanish and English reflects the agglutinative morphological nature of the Basque language.




