Chinese Classical Poetry Matching Dataset (CCPM)
收藏资源简介:
中国古典诗歌匹配数据集(CCPM)是由清华大学计算机科学与技术系的研究团队创建的,旨在通过诗歌匹配任务评估模型对古典中文诗歌的语义理解能力。该数据集包含27,218对中英双语平行数据,涵盖了古典诗歌及其现代中文翻译。创建过程中,研究团队首先从网络收集了6,000段古汉语及其对应的现代中文翻译,然后通过特定的格式过滤和行分割处理,确保数据集的质量和适用性。CCPM数据集的应用领域主要集中在自动分析和生成模型对古典中文诗歌的语义理解,有助于提升相关技术在诗歌创作和分析中的应用。
The Chinese Classical Poetry Matching (CCPM) dataset was developed by a research team from the Department of Computer Science and Technology, Tsinghua University, aimed at evaluating models' semantic understanding of classical Chinese poetry through poetry matching tasks. This dataset contains 27,218 pairs of Chinese-English bilingual parallel data, covering classical Chinese poems and their modern Chinese translations. During the construction of CCPM, the research team first collected 6,000 pieces of classical Chinese poetry and their corresponding modern Chinese translations from the internet, then performed specific format filtering and line segmentation processing to ensure the dataset's quality and applicability. The main application scenarios of the CCPM dataset focus on evaluating the semantic comprehension of classical Chinese poetry by automatic analysis and generation models, which helps to advance the application of related technologies in poetry creation and analysis.

- 1CCPM: A Chinese Classical Poetry Matching Dataset清华大学 · 2021年



