LUXALIGN
收藏资源简介:
LUXALIGN数据集由卢森堡大学创建,旨在为卢森堡语提供高质量的跨语言平行数据。该数据集通过收集RTL.lu新闻平台上的新闻文章,利用OpenAI的文本嵌入模型进行语言对齐,提取了卢森堡语与英语、法语的平行句子。数据集包含25996条卢森堡语-英语和86293条卢森堡语-法语的平行数据。创建过程涉及预处理、过滤和句子对齐,旨在提升低资源语言的句子嵌入模型性能,特别是在信息检索和文档聚类等领域。
The LUXALIGN dataset was developed by the University of Luxembourg to provide high-quality cross-lingual parallel data for Luxembourgish. This dataset collects news articles from the RTL.lu news platform, leverages OpenAI's text embedding models for language alignment, and extracts parallel sentence pairs between Luxembourgish and English as well as French. The dataset includes 25,996 Luxembourgish-English and 86,293 Luxembourgish-French parallel data pairs. Its creation workflow involves preprocessing, filtering and sentence alignment, with the goal of enhancing the performance of sentence embedding models for low-resource languages, especially in domains such as information retrieval and document clustering.




