Diplomatic Chinese-English Parallel Dataset
收藏资源简介:
该数据集由广东工业大学和香港中文大学的研究团队构建,包含5528条高质量的中英平行句子,主要涉及中国外交部发言人答记者问的内容。数据集具有高度的语义一致性,经过严格的校对和审核,适用于神经机器翻译任务。数据集的创建旨在评估自适应少样本提示框架(AFSP)在最新语言上的有效性,并扩展神经机器翻译的研究边界。该数据集的应用领域主要集中在机器翻译,特别是外交领域的文本翻译,旨在提高翻译的语义一致性和准确性。
This dataset was developed by a research team from Guangdong University of Technology and The Chinese University of Hong Kong. It consists of 5,528 high-quality Chinese-English parallel sentences, primarily covering content from press briefings where spokespersons of the Ministry of Foreign Affairs of the People's Republic of China responded to reporters' questions. The dataset features high semantic consistency, has undergone strict proofreading and review, and is suitable for neural machine translation tasks. The purpose of creating this dataset is to evaluate the effectiveness of the Adaptive Few-shot Prompting (AFSP) framework in state-of-the-art neural machine translation applications, and to push forward the research frontier of neural machine translation. Its application fields mainly focus on machine translation, especially text translation in the diplomatic domain, with the goal of improving the semantic consistency and accuracy of translation results.

- 1Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models广东工业大学, 香港中文大学 · 2025年



