CSD-Dataset
收藏资源简介:
CSD-Dataset是由南京大学创建的中文数据集,用于子文本识别研究。该数据集从流行的社交媒体平台如微博、知乎、网易云音乐和哔哩哔哩收集了约70,000条评论数据,经过匿名化处理以保护用户隐私。数据集包含详细的标注信息,包括讽刺、隐喻、夸张等七种信息,通过两轮标注确保质量。CSD-Dataset旨在通过机器学习方法帮助计算机理解文本中的隐含意义,特别是在情感分析和隐喻识别等领域中具有重要应用价值。
CSD-Dataset is a Chinese-language dataset created by Nanjing University for subtext recognition research. It collects approximately 70,000 pieces of comment data from popular social media platforms including Weibo, Zhihu, NetEase Cloud Music and Bilibili, and has been anonymized to protect user privacy. The dataset contains detailed annotation information covering seven categories such as sarcasm, metaphor, hyperbole and others, with its quality ensured through two rounds of annotation. CSD-Dataset aims to help computers understand the implicit meanings in texts via machine learning methods, and has important application value in fields such as sentiment analysis and metaphor recognition.

- 1SASICM A Multi-Task Benchmark For Subtext Recognition南京大学 · 2021年



