C-MTCSD
收藏资源简介:
C-MTCSD是一个中文多轮对话立场检测数据集,由深圳技术大学创建。该数据集从新浪微博收集了24,264个经过精心注释的实例,是迄今为止最大的中文对话立场检测数据集,比之前唯一的中文对话立场检测数据集CANT-CSD大4.2倍。数据集涵盖了科技领域的话题(如iPhone 15、Apollo Go)以及具有争议性的社会话题(如不婚主义、裸辞、预制菜)。数据集的构建经历了数据收集、预处理、注释和质量保证等步骤,最终形成了高质量、多样化的对话语料库。C-MTCSD旨在解决中文立场检测研究中存在的挑战,为相关研究提供了新的基准。
C-MTCSD is a Chinese multi-turn dialogue stance detection dataset developed by Shenzhen Technology University. This dataset includes 24,264 carefully annotated instances collected from Sina Weibo, standing as the largest Chinese dialogue stance detection dataset to date, which is 4.2 times larger than the only prior Chinese dialogue stance detection dataset CANT-CSD. It covers topics in the field of science and technology (e.g., iPhone 15, Apollo Go) and controversial social issues, including non-maritalism, unannounced resignation, and pre-made meals. The development of the dataset involved procedures such as data collection, preprocessing, annotation, and quality assurance, ultimately resulting in a high-quality and diverse dialogue corpus. C-MTCSD aims to address the existing challenges in Chinese stance detection research, offering a novel benchmark for relevant studies.




