Dream of the Red Chamber corpus
收藏资源简介:
该数据集名为“Dream of the Red Chamber corpus”,由上海交通大学与华东理工大学的研究团队联合构建,旨在系统探究文化负载翻译在机器翻译中的挑战。数据集内容基于中国古典文化代表作《红楼梦》构建,包含500个精心筛选的中日双语文化负载片段,均匀涵盖生态、宗教、物质、语言和社会五大文化类别,源文本平均长度约24词,目标文本约51词,数据来源于伊藤漱平翻译的权威中日双语版本。创建过程涉及四名中文研究生初步标注与四名日语研究生二次筛选,耗时20天完成分类与平衡处理。该数据集主要应用于跨文化机器翻译研究领域,着力解决大语言模型在文化负载表达翻译中的性能瓶颈,并为文化导向的翻译评估提供基准。
The dataset named "Dream of the Red Chamber corpus" was jointly constructed by the research teams from Shanghai Jiao Tong University and East China University of Science and Technology, with the goal of systematically investigating the challenges of culturally loaded translation in machine translation. The content of this dataset is constructed based on *Dream of the Red Chamber*, a representative masterpiece of classical Chinese culture, and contains 500 carefully selected Chinese-Japanese bilingual culturally loaded segments that evenly cover five major cultural categories: ecology, religion, material culture, language, and society. The average length of the source texts is approximately 24 words, while that of the target texts is about 51 words. The dataset is sourced from the authoritative Chinese-Japanese bilingual edition translated by Ito Shūhei. The construction process involved four graduate students majoring in Chinese conducting preliminary annotation, followed by secondary screening by four graduate students majoring in Japanese, and the classification and balancing work was completed within 20 days. This dataset is primarily applied in the field of cross-cultural machine translation research, aiming to address the performance bottlenecks of large language models (LLMs) in translating culturally loaded expressions, and to serve as a benchmark for culture-oriented translation evaluation.

- 1On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens上海交通大学·计算机学院; 华东理工大学·外语学院 · 2026年




