Public Sentiment Dataset of Bangladeshi People Towards Government in Social Media
收藏资源简介:
This dataset contains 6,000 manually annotated Bangla Facebook comments representing public sentiment toward the Interim Government of Bangladesh, collected between August 2024 and 2025. The collection period falls immediately following the July 2024 political uprising that led to a major governmental transition in the country. The central research hypothesis is that Bangla political sentiment expressed on social media can be reliably classified into three categories (Positive, Negative, and Neutral), and that language-specific transformer models can capture the nuanced, context-dependent nature of such discourse more effectively than traditional machine learning approaches. The dataset is perfectly balanced, with exactly 2,000 comments per sentiment class. Comments were collected from major public Facebook pages including Chief Advisor GOB, Jamuna Television, Somoy News, Daily Prothom Alo, Kalbela Digital, and Dhaka Tribune, ensuring diversity across both news outlets and political discussion spaces. "Positive" comments tend to express approval, trust, or satisfaction toward government actions; "Negative" comments reflect criticism, dissatisfaction, or distrust; and "Neutral" comments include factual statements, questions, or ambiguous opinions without clear polarity. Comments were collected using a hybrid approach combining manual selection, browser extension-based extraction, and automated scraping via Apify. Only publicly accessible posts and comments were included; private profiles, closed groups, and personally identifiable information were excluded. Three annotators with Bengali language proficiency and social media political awareness labeled each comment independently into one of three sentiment classes (Positive, Negative, Neutral) following a detailed annotation guideline. An expert reviewer with 15 years of experience resolved all disagreements and validated borderline cases. Inter-annotator reliability was measured using Fleiss' Kappa, yielding a score of κ = 0.91, indicating almost perfect agreement and confirming the reliability of the annotations. The dataset is provided in two files. "raw_data_bangla_public_sentiment_toward_government.xlsx" contains the original annotated comments with two columns: Sentence (raw Bangla text) and Sentiment (label: Positive, Negative, or Neutral). "preprocessed_data_bangla_public_sentiment_toward_government.xlsx" contains three columns: Sentence (original text), Sentiment (label), and Clean Sentence (preprocessed text after Unicode normalization, removal of HTML tags, URLs, mentions, hashtags, digits, emojis, non-Bangla characters, and extra whitespace). Researchers may use the raw file for custom preprocessing pipelines or the preprocessed file directly for feature extraction and model training. The dataset is suitable for Bangla sentiment analysis, NLP benchmarking, political discourse analysis, and training or fine-tuning transformer-based models for low-resource Bangla text classification.
本数据集包含6000条经人工标注的孟加拉语Facebook评论,用于反映民众对孟加拉国临时政府的公众情感倾向,采集时段为2024年8月至2025年。该采集时段紧随2024年7月引发该国重大政府更迭的政治动荡之后。 本研究的核心假说为:社交媒体上表达的孟加拉语政治情感可被可靠划分为三类(积极、消极与中性);针对特定语言的Transformer模型(Transformer)相较于传统机器学习方法,能更有效地捕捉此类话语中细微且依赖上下文的语义特征。 本数据集分布完全均衡,每个情感类别恰好包含2000条评论。评论采集自多个主流公共Facebook页面,包括孟加拉国首席顾问公署(Chief Advisor GOB)、朱纳电视台(Jamuna Television)、索莫伊新闻(Somoy News)、《每日普罗坦阿洛报》(Daily Prothom Alo)、卡尔贝拉数字媒体(Kalbela Digital)及《达卡论坛报》(Dhaka Tribune),确保覆盖新闻媒体与政治讨论空间的多样性。 “积极”评论通常表达对政府举措的认可、信任或满意;“消极”评论体现为批评、不满或不信任;“中性”评论则包含事实陈述、疑问或无明确情感倾向的模糊观点。 评论采集采用混合方案,结合人工筛选、基于浏览器扩展的提取以及通过Apify平台的自动爬取。仅纳入公开可见的帖子与评论,排除私人账号、封闭群组及个人可识别信息。三名具备孟加拉语能力且熟悉社交媒体政治议题的标注人员,依据详细的标注指南,独立将每条评论划分为上述三类情感标签。一名拥有15年工作经验的专家审核员负责解决所有标注分歧,并对边界案例进行验证。采用弗莱施Kappa系数(Fleiss' Kappa)衡量标注者间信度,最终得分为κ=0.91,表明标注一致性近乎完美,证实了标注结果的可靠性。 本数据集以两个文件形式提供。"raw_data_bangla_public_sentiment_toward_government.xlsx"包含原始标注评论,共两列:`Sentence`(原始孟加拉语文本)与`Sentiment`(情感标签:积极、消极或中性)。"preprocessed_data_bangla_public_sentiment_toward_government.xlsx"包含三列:`Sentence`(原始文本)、`Sentiment`(情感标签)与`Clean Sentence`(经预处理后的清洁文本,已完成Unicode归一化、移除HTML标签、统一资源定位符(URL)、提及、话题标签、数字、表情符号、非孟加拉语字符及多余空格)。 研究人员可使用原始文件进行自定义预处理流程,或直接使用预处理文件开展特征提取与模型训练。本数据集适用于孟加拉语情感分析、自然语言处理(Natural Language Processing,简称NLP)基准测试、政治话语分析,以及针对低资源孟加拉语文本分类任务的基于Transformer的模型训练与微调。



