FinanceInc/auditor_sentiment
收藏资源简介:
Auditor Sentiment数据集是一个用于情感分类的数据集,包含数千条来自英文财经新闻的句子,每条句子都被标注为正面、中性或负面情感。数据集的创建目的是为了提高情感分类的准确性,之前的现成情感分类工具只能达到70%的F1分数。数据集由16位具有金融市场背景知识的专家进行标注,且标注一致性超过75%。数据集的结构包括句子和对应的情感标签,数据被随机分为训练集和测试集,比例为75/25。数据集的语言为英语,且不包含个人或敏感信息。
Auditor Sentiment Dataset is a sentiment classification dataset containing thousands of sentences sourced from English financial news. Each sentence is annotated with one of three sentiment labels: positive, neutral, or negative. This dataset was developed to improve the accuracy of sentiment classification, as existing off-the-shelf sentiment classification tools only achieved a 70% F1-score. It was annotated by 16 experts with financial market background knowledge, with inter-annotator agreement exceeding 75%. The dataset structure includes sentences and their corresponding sentiment labels, and the data is randomly split into training and test sets at a 75:25 ratio. The dataset is in English and does not contain any personal or sensitive information.
数据集概述
数据集名称
- 名称: Auditor_Sentiment
- 别名: Auditor Sentiment
数据集描述
- 描述: 该数据集包含从金融新闻中提取的数千个英文句子,按情感进行分类。
- 目的: 收集审计员评价,以提高情感分析的性能。
语言和多语言性
- 语言: 英语
- 多语言性: 单语种
数据集大小和类别
- 大小: 1K<n<10K
任务和支持的任务
- 任务: 文本分类
- 支持的任务: 多类分类, 情感分类
数据集结构
- 数据实例: 每个实例包含一个句子及其对应的情感标签(positive, neutral, negative)。
- 数据字段:
- sentence: 数据集中的一个分词行
- label: 对应的类别标签,字符串形式:positive - (2), neutral - (1), negative - (0)
- 数据分割: 随机创建的训练/测试分割,比例为75/25。
数据集创建
- 来源数据: 英文新闻报告
- 注释过程: 由16名具有金融市场背景知识的人员对4840个句子进行注释,选择内部注释一致性大于75%的子集。
- 注释者: 来自SME列表,具体姓名由sue@demo.org持有。
使用数据注意事项
- 偏见讨论: 所有注释者来自同一机构,因此在理解内部注释一致性时应考虑此因素。
- 许可证: Demo.Org Proprietary - DO NOT SHARE




