遇见数据集

BSC-LT/Legal_Catalan_Spanish_Parallel_Corpus

收藏
Hugging Face2026-05-21 更新2026-06-14 收录
官方服务:

资源简介:

法律加泰罗尼亚语-西班牙语平行语料库是一个多语言数据集,包含从加泰罗尼亚公共机构官方文件中提取的法律领域真实平行文本,涉及加泰罗尼亚语和西班牙语。它包含三个子语料库,按不同文本粒度组织:句子级、段落级和文档级。这种多级结构支持超越传统句子级机器翻译的研究和训练,有助于开发能够处理更长文本并在更长文本跨度上产生更连贯翻译的系统。

The Legal Catalan-Spanish Parallel Corpus is a multilingual dataset containing authentic parallel legal texts in Catalan and Spanish, extracted from official documents of Catalan public institutions. It comprises three sub-corpora organized by different text granularities: sentence-level, paragraph-level, and document-level. This multi-level structure supports research and training beyond traditional sentence-level machine translation, aiding the development of systems capable of handling longer texts and generating more coherent translations across longer text spans.

提供机构:
BSC-LT
二维码
社区交流群
二维码
科研交流群
商业服务