SFU Opinion and Comments Corpus
收藏资源简介:
The SFU Opinion and Comments Corpus (SOCC) is a corpus for the analysis of online news comments. Our corpus contains comments and the articles from which the comments originated. The articles are all opinion articles, not hard news articles. The corpus is larger than any other currently available comments corpora, and has been collected with attention to preserving reply structures and other metadata. In addition to the raw corpus, we also present annotations for four different phenomena: constructiveness, toxicity, negation and its scope, and appraisal. The data is divided into two main parts: raw data and annotated data. The raw data contains three CSVs: gnm_artcles.csv, gnm_comments.csv, and gnm_comment_threads.csv. The annotated data contains annotations for constructiveness, negation, and appraisal. The details of our different corpora and how to use them are on the following GitHub page. https://github.com/sfu-discourse-lab/SOCC/blob/master/README.md
SFU观点与评论语料库(SFU Opinion and Comments Corpus, SOCC)是一款用于在线新闻评论分析的专业语料库。本语料库收录了评论及其所属的源文章,所有源文章均为观点类文章,而非硬新闻稿件。该语料库规模大于当前所有可获取的评论类语料库,且在采集过程中注重保留评论回复结构及其他元数据。除原始语料库外,本研究还提供了四类不同语言现象的标注:评论建设性、评论毒性、否定结构及其辖域,以及评价标注。数据集分为两大模块:原始数据与标注数据。其中原始数据包含三个CSV文件:gnm_artcles.csv、gnm_comments.csv 与 gnm_comment_threads.csv;标注数据则包含建设性、否定结构及评价三类标注。本语料库的详细说明与使用方法详见以下GitHub页面:https://github.com/sfu-discourse-lab/SOCC/blob/master/README.md




