sungkwang2/klingon-en-tlh-translation
收藏资源简介:
该数据集是一个用于英语和克林贡语(Klingon)双向机器翻译的数据集,基于OPUS Tatoeba en-tlh v2023-04-12(采用CC-BY 4.0许可)构建,并经过重复数据删除处理。数据集包含训练集(12,717条样本)、验证集(500条样本)和测试集(500条样本)。数据格式为JSONL,每个文件对应不同的翻译方向:*.jsonl文件用于英语→克林贡语翻译,*_tlh2en.jsonl文件用于克林贡语→英语翻译。每条数据包含instruction、output、en(英语文本)和klingon(克林贡语文本)字段。该数据集旨在支持自然语言处理研究,特别是用于比较自回归模型(如Qwen2.5-7B)和扩散模型(如Fast-dLLM v2 7B)在翻译任务上的性能。
This dataset is a bilingual machine translation dataset for English and Klingon, based on OPUS Tatoeba en-tlh v2023-04-12 (under CC-BY 4.0 license) with deduplication applied. It includes a training set (12,717 samples), validation set (500 samples), and test set (500 samples). The data is in JSONL format, with separate files for translation directions: *.jsonl for English→Klingon translation and *_tlh2en.jsonl for Klingon→English translation. Each line contains fields: instruction, output, en (English text), and klingon (Klingon text). The dataset is designed for NLP research, specifically for comparing the performance of autoregressive models (e.g., Qwen2.5-7B) and diffusion models (e.g., Fast-dLLM v2 7B) on translation tasks.




