遇见数据集

Four Text Datasets Used For Comparison Between Hedonometer and Azure Sentiment Analysis Tools

收藏
Figshare2023-11-09 更新2026-04-08 收录
官方服务:

资源简介:

Lexicon-based approaches to sentiment analysis of text are based on each word or lexical entry having a pre-defined<br>weight indicating its sentiment polarity. We compute sentiment for more than 150,000 English language texts drawn from 4 domains using the Hedonometer, a lexicon-based technique and Azure, a contemporary machine-learning based approach. We model differences in sentiment scores between approaches for documents in each domain using a regression and analyse the independent variables (Hedonometer lexical entries) as indicators of each word's importance and contribution to the score differences.1. Finance Data: This dataset contains 5,000 records of different financial news texts from company press reviews and news headlines.2. News Headlines Data: This dataset consists of 50,000 news headlines for the period of 8 months (November 2015 to July 2016) on four different topics: Economy, Microsoft, Obama, and Palestine.3. IMDb Dataset: This dataset consists of 50,000 reviews posted by customers on the online IMDb platform which is an International Movie Database platform.4. Twitter Dataset: This dataset consists of almost 40,000 tweets from users around the globe on every thing.5. Hedonometer Bag of Words: This is the bag of words used to perform sentiment analysis using traditional lexicon approach which consists of 10,223 words with their respective happiness score. The actual file can be downloaded from here: https://hedonometer.org/words/labMT-en-v2/6. Combined p-values results: This is the result file which was generated once we performed sentiment analysis on all the above domains and only identified words that are present in the hedonometer sheet. The sheet consists of the words and their respective happiness score and their p-values on all different domains.7. Data visualisations: This is the visualisation code base in Tableau which was used to generate visualisations.

基于词典的文本情感分析方法,依托每个单词或词项所附带的预定义权重,以表征其情感极性。本研究采用基于词典的情感分析工具Hedonometer,以及当前主流的机器学习方法Azure,对来自4个领域的15万余篇英文文本进行情感计算。我们针对每个领域的文档,通过回归模型刻画不同分析方法间的情感评分差异,并以Hedonometer词项作为自变量,分析其对评分差异的重要性与贡献度。 1. 金融数据集:本数据集包含5000条来自企业新闻稿与财经新闻标题的金融文本记录。 2. 新闻标题数据集:本数据集涵盖2015年11月至2016年7月共8个月内的50000条新闻标题,涉及经济、Microsoft、奥巴马与巴勒斯坦四大主题。 3. IMDb(Internet Movie Database)数据集:本数据集包含50000条来自国际电影数据库IMDb平台的用户影评。 4. Twitter数据集:本数据集包含来自全球用户的近40000条各类主题推文。 5. Hedonometer词袋:该词袋为采用传统词典法开展情感分析所用的词表,包含10223个单词及其对应的快乐评分。该词表文件可从以下链接下载:https://hedonometer.org/words/labMT-en-v2/ 6. 组合p值结果文件:该文件为我们对上述所有领域文本开展情感分析后生成,仅包含出现在Hedonometer词表中的单词,表中列示各单词及其对应快乐评分与各领域下的p值。 7. 可视化代码库:该代码基于Tableau开发,用于生成本次研究的可视化成果。

创建时间:
2023-11-09
二维码
社区交流群
二维码
科研交流群
商业服务