价值观量化评价数据集
收藏资源简介:
价值观量化评价数据集主要面向主流价值观量化分析与模型训练研究需求建设,基于权威媒体和自媒体平台的新闻文本及短文本数据产生。 价值观量化评价数据根据数据的来源和用途可以分为3个子数据集,分别为:新闻文本共合计2.6万条数据,12万条句子粒度的短文本数据,用于量化讽刺类型的约8千条数据。 新闻文本数据集用于预测整体发帖文本(长度约几百字)的整体价值观的量化评估,包括预测文本主要涉及的价值观(七分类)以及价值观的正负极性(三分类),数据集来源为新浪、腾讯、澎湃三个国内新闻网站的新闻;微博、B站等平台上的自媒体。 句子粒度数据集用于预测新闻中单句(长度约十几到几十字不等)的价值观的量化评估,同样也包括句子涉及到的主要价值观(七分类)和价值观的正负极性(三分类),数据来源为融合量化整体新闻文本拆解而来。 量化讽刺数据集用于预测新闻评论区下用户评论(十几字)的讽刺识别,包括预测讽刺类别(四分类)以及是否为讽刺评论(二分类),数据来源主要为2023年1月至11月期间新浪微博下热点微博和利用ChatGPT相应扩增构造。
The Values Quantification Evaluation Dataset is constructed to meet the research requirements of mainstream values quantification analysis and model training, and is generated from news texts and short text data sourced from authoritative media and self-media platforms. The dataset is divided into three subsets based on data sources and application scenarios: 26,000 news text entries, 120,000 sentence-level short text entries, and approximately 8,000 entries for sarcasm type quantification. The news text subset is designed for quantitative evaluation of the overall values of full published texts (typically hundreds of characters in length), which involves predicting the core values covered by the text (7-class classification) and the polarity of these values (3-class classification). The data is sourced from news published on three domestic news websites: Sina, Tencent, and The Paper, as well as self-media content from platforms including Weibo and Bilibili. The sentence-level subset is used for quantitative evaluation of the values of individual sentences within news articles (ranging from ten to several dozen characters in length). It also covers two prediction tasks: identifying the core values involved in the sentence (7-class classification) and determining the polarity of these values (3-class classification). This subset's data is generated by decomposing full news texts and performing quantification processing. The sarcasm quantification subset is intended for sarcasm recognition in user comments (around ten to several dozen characters) from news comment sections. Its tasks include predicting the sarcasm category (4-class classification) and judging whether a comment is sarcastic (2-class classification). The data primarily comes from hot Weibo posts on Sina Weibo between January and November 2023, as well as content augmented and constructed using ChatGPT.




