SEN - Sentiment analysis of Entities in News headlines
收藏资源简介:
If you wish to use this data please cite: Katarzyna Baraniak, Marcin Sydow, A dataset for Sentiment analysis of Entities in News headlines (SEN), Procedia Computer Science, Volume 192, 2021, Pages 3627-3636, ISSN 1877-0509, https://doi.org/10.1016/j.procs.2021.09.136. (https://www.sciencedirect.com/science/article/pii/S1877050921018755) bibtex: users.pja.edu.pl/~msyd/bibtex/sydow-baraniak-SENdataset-kes21.bib SEN is a novel publicly available human-labelled dataset for training and testing machine learning algorithms for the problem of entity level sentiment analysis of political news headlines. On-line news portals play a very important role in the information society. Fair media should present reliable and objective information. In practice there is an observable positive or negative bias concerning named entities (e.g. politicians) mentioned in the on-line news headlines. Our dataset consists of 3819 human-labelled political news headlines coming from several major on-line media outlets in English and Polish. Each record contains a news headline, a named entity mentioned in the headline and a human annotated label (one of “positive”, “neutral”, “negative” ). Our SEN dataset package consists of 2 parts: SEN-en (English headlines that split into SEN-en-R and SEN-en-AMT), and SEN-pl (Polish headlines). Each headline-entity pair was annotated via team of volunteer researchers (the whole SEN-pl dataset and a subset of 1271 English records: the SEN-en-R subset, “R” for “researchers”) or via the Amazon Mechanical Turk service (a subset of 1360 English records: the SEN-en-AMT subset). During analysis of annotation outlying annotations and removed . Separate version of dataset without outliers is marked by "noutliers" in data file name. Details of the process of preparing the dataset and presenting its analysis are presented in the paper. In case of any questions, please contact one of the authors. Email adresses are in the paper.
若您使用本数据集,请引用如下文献:Katarzyna Baraniak、Marcin Sydow,《新闻标题中实体的情感分析数据集(SEN)》,《Procedia Computer Science》(计算机科学进展),2021年,第192卷,第3627-3636页,ISSN 1877-0509,https://doi.org/10.1016/j.procs.2021.09.136(https://www.sciencedirect.com/science/article/pii/S1877050921018755)。BibTeX引用文件:users.pja.edu.pl/~msyd/bibtex/sydow-baraniak-SENdataset-kes21.bib。 SEN是一款全新的公开可用人工标注数据集,用于训练与测试针对政治新闻标题的实体级情感分析(entity level sentiment analysis)机器学习算法。在线新闻门户在信息社会中发挥着至关重要的作用。公正媒体应呈现可靠且客观的信息,但实际场景中,在线新闻标题提及的命名实体(named entity,如政治人物)往往存在可观测的正向或负向偏见。 本数据集包含来自多家主流英文与波兰语在线媒体的3819条经人工标注的政治新闻标题。每条记录均包含一则新闻标题、标题中提及的命名实体,以及人工标注的情感标签(分为“正向(positive)”、“中性(neutral)”、“负向(negative)”三类)。 本SEN数据集套件包含两个部分:SEN-en(英语新闻标题,进一步划分为SEN-en-R与SEN-en-AMT)以及SEN-pl(波兰语新闻标题)。每条标题-实体对的标注工作由志愿者研究团队完成(覆盖全部SEN-pl数据集与1271条英语记录子集,即SEN-en-R子集,其中“R”代表“研究者(researchers)”),或通过亚马逊机械Turk(Amazon Mechanical Turk)服务完成(对应1360条英语记录子集,即SEN-en-AMT子集)。 在对标注结果进行分析的过程中,异常标注已被移除。数据集的无异常值版本会在数据文件名中以“noutliers”标识。本数据集的制备流程与分析细节已在上述论文中详述。若有任何疑问,请联系其中一位作者,其电子邮箱地址可在论文中获取。




