遇见数据集

ABSA Beer dataset and code

收藏
Zenodo2026-05-04 更新2026-05-26 收录
官方服务:

资源简介:

The datasets are results of the paper Unsupervised Aspect-Based Sentiment Analysis through LLM: A Case Study of an Unlabeled Portuguese Beer Database Data sets: Reviews Main (step_3_reviews_main.csv): Step 3 produced this dataset, which supports the analysis of review comments, quantitative beer attributes, and review-related information, as summarized in Table 2. From the 67,083 reviews obtained in the previous step, 7,100 reviews (10.58%) were classified as non-relevant, while 59,982 reviews (89.42%) were deemed relevant. Consequently, the final dataset comprises 59,982 records. Reviews Sample (step_4_1_reviews_sample.csv): Step 4 generated this dataset as a subset of the Reviews Main, with 108 reviews, particularly suitable for benchmarking and validating new ABSA methodologies along side the ABSA Gold base. ABSA Gold (step_4_ABSA_Gold.csv): This manually annotated dataset, derived from Reviews Sample, serves as a gold standard for comparative evaluation of prompt-based approaches applied to the Reviews Sample dataset. It contains 1710 annotated BC from 108 reviews (columns: index, aspect, category, sentiment). ABSA Main (step_4_ABSA_main.csv): The final dataset produced in Step 4 is designed to enable large-scale analysis and knowledge discovery within the beer industry. It comprises 880,373 records and supports the extraction of novel insights from consumer reviews (columns: index, aspect, category, sentiment). ABSA Final (step_5_ABSA-FINAL.csv): The dataset produced in Step 5 comprises only essential information for extracting the results from this step. Columns: index, review_comment, review_datetime, beer_style, review_general_rate, review_aroma, review_visual, review_flavor, review_sensation, review_general_set, aspect, category, sentiment, year.

本数据集均来自论文《基于大语言模型(LLM)的无监督面向方面情感分析(Aspect-Based Sentiment Analysis,ABSA):葡萄牙未标注啤酒数据库的案例研究》。 数据集详情如下: 主评论数据集(step_3_reviews_main.csv): 该数据集由步骤3生成,可用于分析评论内容、量化啤酒属性及评论相关信息,具体信息汇总于表2。从前一步获取的67083条评论中,7100条(占比10.58%)被归类为无关评论,剩余59982条(占比89.42%)被认定为有效评论,因此本数据集最终包含59982条数据记录。 采样评论数据集(step_4_1_reviews_sample.csv): 该数据集由步骤4生成,为主评论数据集的子集,共包含108条评论,尤其适用于结合ABSA金标准数据集对新型ABSA方法开展基准测试与验证。 ABSA金标准数据集(step_4_ABSA_Gold.csv): 该数据集为从采样评论数据集中提取的人工标注数据集,可作为面向采样评论数据集的提示词方法对比评估的金标准。其包含来自108条评论的1710条标注数据,字段包括:索引、方面、类别、情感倾向。 主ABSA数据集(step_4_ABSA_main.csv): 该数据集为步骤4生成的最终数据集,旨在支持啤酒行业内的大规模分析与知识发现。数据集共包含880373条数据记录,可用于从消费者评论中提取新颖见解,字段包括:索引、方面、类别、情感倾向。 最终ABSA数据集(step_5_ABSA-FINAL.csv): 该数据集由步骤5生成,仅包含用于提取本步骤结果的必要信息。字段包括:索引、评论内容、评论时间、啤酒风格、综合评分、香气、外观、风味、口感、review_general_set、方面、类别、情感倾向、年份。

提供机构:
Zenodo
创建时间:
2026-02-12
二维码
社区交流群
二维码
科研交流群
商业服务