CMDAG
收藏资源简介:
CMDAG是一个由香港科技大学等机构合作创建的大型中文隐喻数据集,包含约28,000条从多种中文文学来源(如诗歌、散文、歌词等)提取的句子。该数据集特别之处在于每条隐喻句子都附带有其对应的‘喻意’(GROUNDS)。创建过程中,研究团队利用了专业的标注者进行精细标注,确保了数据的质量和一致性。CMDAG数据集主要用于支持中文隐喻生成的研究,特别是在机器学习和自然语言处理领域,旨在提高模型生成隐喻句子的创造性和流畅性。
CMDAG is a large-scale Chinese metaphor dataset co-developed by institutions including the Hong Kong University of Science and Technology, containing approximately 28,000 sentences extracted from diverse Chinese literary sources such as poetry, prose, and lyrics. A key characteristic of this dataset is that each metaphorical sentence is accompanied by its corresponding GROUNDS, a term referring to the underlying figurative meaning of the metaphor. During the dataset construction process, the research team engaged professional annotators to perform fine-grained annotation, which guarantees the high quality and consistency of the data. The CMDAG dataset is primarily designed to support research on Chinese metaphor generation, especially within the domains of machine learning and natural language processing, with the objective of improving the creativity and fluency of metaphorical sentence generation by AI models.




