LanguageWave (LW)
收藏资源简介:
LanguageWave (LW) 数据集是一个针对黎巴嫩方言的文化感知英黎平行数据集,包含约 3000 个句子,由 95 期播客节目提取而来,涵盖了黎巴嫩文化的各个方面。该数据集由美国贝鲁特美国大学电气与计算机工程系的研究人员创建,旨在解决低资源方言翻译的挑战,特别是黎巴嫩方言的翻译。数据集的创建过程采用了合成数据的方法,利用了黎巴嫩语法书中的规则和示例,并通过 Claude 3.5 Sonnet 生成了相关的翻译示例。该数据集已被用于训练和评估大型语言模型在黎巴嫩方言翻译任务上的性能,并取得了优于非原生翻译数据集的成果。数据集的访问地址为 https://github.com/silvayakhni/language_wave。
The LanguageWave (LW) dataset is a culturally-aware English-Lebanese parallel dataset containing approximately 3,000 sentences extracted from 95 podcast episodes, covering various aspects of Lebanese culture. This dataset was developed by researchers from the Department of Electrical and Computer Engineering at the American University of Beirut, aiming to address the challenges of low-resource dialect translation, especially for Lebanese Arabic. The dataset construction adopted a synthetic data generation method, leveraging rules and examples from Lebanese Arabic grammar books and generating corresponding translation examples via Claude 3.5 Sonnet. It has been used to train and evaluate the performance of large language models on Lebanese Arabic translation tasks, and achieved better results than non-native translation datasets. The dataset can be accessed at https://github.com/silvayakhni/language_wave.




