Motamot
收藏资源简介:
Motamot数据集由阿赫桑尼科技大学的计算机科学与工程系创建,专门用于孟加拉语的政治情感分析。该数据集包含7,058条标注了正面和负面情感的数据,来源于多个在线新闻门户,为政治情感分析提供了全面的资源。数据集的创建过程包括从多个新闻源抓取文章和意见文章,并进行细致的手动标注。该数据集主要应用于孟加拉国选举期间的公众意见分析,旨在帮助理解选民偏好和当前趋势。
The Motamot Dataset was created by the Department of Computer Science and Engineering at Ahsanullah University of Science and Technology, specifically for Bengali political sentiment analysis. It contains 7,058 labeled samples with positive and negative sentiment annotations, sourced from multiple online news portals, thus serving as a comprehensive resource for political sentiment analysis research. The dataset development process involves scraping articles and opinion pieces from various news sources, followed by meticulous manual annotation. This dataset is primarily applied to public opinion analysis during elections in Bangladesh, with the aim of helping researchers understand voter preferences and prevailing societal trends.
Motamot 数据集概述
数据集介绍
Motamot 数据集是一个用于分析孟加拉语政治情感的数据集,从多个在线新闻报纸中精心收集,涵盖了孟加拉国选举期间的政治事件和对话。数据集包括文章和观点文章,确保了政治话语的多样性和代表性。
数据集结构
数据分割
| 数据类型 | 训练集 | 测试集 | 验证集 |
|---|---|---|---|
| 总数 | 5647 | 706 | 705 |
| 正面情感 | 3306 | 413 | 413 |
| 负面情感 | 2341 | 293 | 292 |
训练数据
| 情感类型 | 数量 |
|---|---|
| 正面 | 3306 |
| 负面 | 2341 |
测试数据
| 情感类型 | 数量 |
|---|---|
| 正面 | 413 |
| 负面 | 293 |
验证数据
| 情感类型 | 数量 |
|---|---|
| 正面 | 413 |
| 负面 | 292 |
性能比较
预训练语言模型性能
| 模型 | 准确率 | 精确率 | 召回率 | F1分数 |
|---|---|---|---|---|
| BanglaBERT | 0.8204 | 0.8222 | 0.8204 | 0.8203 |
| Bangla BERT Base | 0.6803 | 0.6907 | 0.6812 | 0.6833 |
| DistilBERT | 0.6320 | 0.6358 | 0.6320 | 0.6317 |
| mBERT | 0.6427 | 0.6496 | 0.6428 | 0.6153 |
| sahajBERT | 0.6708 | 0.6791 | 0.6709 | 0.6707 |
大型语言模型性能
| 模型 | 指标 | Zero-shot | 5-shot | 10-shot | 15-shot |
|---|---|---|---|---|---|
| GPT 3.5 Turbo | 准确率 | 0.8500 | 0.8900 | 0.9133 | 0.9400 |
| 精确率 | 0.8467 | 0.8867 | 0.9200 | 0.9467 | |
| 召回率 | 0.8533 | 0.8926 | 0.9079 | 0.9342 | |
| F1分数 | 0.8495 | 0.8896 | 0.9139 | 0.9404 | |
| Gemini 1.5 Pro | 准确率 | 0.8608 | 0.8981 | 0.9200 | 0.9633 |
| 精确率 | 0.8931 | 0.8846 | 0.9333 | 0.9667 | |
| 召回率 | 0.8477 | 0.9205 | 0.9091 | 0.9603 | |
| F1分数 | 0.8698 | 0.9022 | 0.9211 | 0.9635 |

- 1Motamot: A Dataset for Revealing the Supremacy of Large Language Models over Transformer Models in Bengali Political Sentiment Analysis计算机科学与工程系 阿赫桑尼科技大学 (AUST), 达卡, 孟加拉国 · 2024年



