EN2CS
收藏资源简介:
EN2CS数据集是由自然发生的英语-西班牙语代码转换句子和对应的英语单语句子组成的伪平行语料库。该数据集由HiTZ Center - Ixa研究机构创建,旨在训练和评估英语-西班牙语代码转换生成模型。数据集包括训练、开发和测试三个部分,共有12933条训练数据。该数据集的创建过程包括从LINCE基准数据集中筛选出代码转换句子,使用Command R模型生成英语单语句子,然后进行后编辑以创建黄金标准测试集。该数据集的应用领域是代码转换文本的生成,旨在解决机器翻译和自然语言处理中的代码转换问题。
The EN2CS dataset is a pseudo-parallel corpus comprising naturally occurring English-Spanish code-switching sentences and their corresponding English monolingual sentences. It was developed by the HiTZ Center - Ixa research institute for training and evaluating English-Spanish code-switching generation models. The dataset is split into three subsets: training, development, and test, with a total of 12,933 training samples. Its construction process includes screening code-switching sentences from the LINCE benchmark dataset, generating English monolingual sentences using the Command R model, and performing post-editing to build the gold-standard test set. This dataset targets code-switching text generation, aiming to address code-switching-related challenges in machine translation and natural language processing.

- 1Conditioning LLMs to Generate Code-Switched Text: A Methodology Grounded in Naturally Occurring DataHiTZ Center - Ixa, University of the Basque Country UPV/EHU · 2025年



