proxectonos/es-gl_museo_virtual_usc_descricions_parallel
收藏资源简介:
`es-gl_museo_virtual_usc_descricions_parallel` 是一个西班牙语-加利西亚语平行语料库,包含与USC虚拟博物馆相关的文化遗产描述对齐文本。该数据集旨在支持机器翻译、领域适应以及在文化遗产和博物馆领域开发加利西亚语资源。数据集由两个对齐的纯文本文件组成:`es.txt`(西班牙语部分)和`gl.txt`(加利西亚语部分),每行对齐,即`es.txt`中的每一行对应`gl.txt`中相同行号的内容。在Hugging Face数据集查看器中,两种语言作为单独配置提供:`es`(西班牙语描述)和`gl`(加利西亚语描述)。数据来源于USC虚拟博物馆的描述和编目信息,经过处理、清洗、归一化和对齐,适用于西班牙语-加利西亚语机器翻译和文化遗产语言技术实验。
`es-gl_museo_virtual_usc_descricions_parallel` is a Spanish-Galician parallel corpus containing aligned cultural heritage descriptions associated with the Museo Virtual da USC. The dataset is intended to support machine translation, domain adaptation, and the development of Galician language resources in the cultural heritage and museum domains. The dataset consists of two aligned plain-text files: `es.txt` (Spanish side) and `gl.txt` (Galician side), with line alignment where each line in `es.txt` corresponds to the same line number in `gl.txt`. In the Hugging Face dataset viewer, the two languages are available as separate configurations: `es` (Spanish descriptions) and `gl` (Galician descriptions). It was created from descriptive and cataloguing information related to the Museo Virtual da USC, processed and organized for Spanish-Galician machine translation and cultural heritage language technology experiments.




