遇见数据集

Penerapan BERT dalam Memetakan Opini Pengguna Instagram Terhadap Program Makan Bergizi Gratis

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This dataset consists of Instagram comments collected from users in South Sumatra discussing the MBG (Free Lunch Program), obtained through hashtag-based scraping using the keyword #MBG. A total of 14,742 comments were collected from 93 Instagram posts published between February 15, 2025, and October 12, 2025. The data collection process was conducted using a Python-based web scraping script utilizing the Instagrapi library executed in the Google Colab environment to retrieve publicly available comments associated with the selected hashtag. To ensure data quality, a validation process was performed by removing empty comment entries and ensuring data completeness. The dataset then underwent preprocessing stages including text cleaning, tokenization using NLTK, stopword removal, and stemming using the Sastrawi library. After preprocessing, the dataset size was reduced from 14,742 comments to 14,051 validated comments, which were used as the final dataset for analysis. This dataset can be utilized for sentiment analysis, public opinion monitoring, policy evaluation, and Natural Language Processing (NLP) research, providing empirical insights into community responses toward public welfare programs on social media. Furthermore, the dataset can be continuously reused for future research, comparative studies, and the development of data-driven policy evaluation models, supporting sustainable and long-term analysis of public sentiment trends.

本数据集源自南苏门答腊地区用户针对MBG(免费午餐计划,Free Lunch Program)的Instagram评论,通过以#MBG为关键词的话题标签爬取方式获取。研究团队于2025年2月15日至2025年10月12日期间,从93条Instagram帖文中共收集到14742条评论。本数据集的采集流程采用基于Python的网络爬虫脚本完成,该脚本借助Instagrapi库,在Google Colab环境中运行,以获取与选定话题标签相关的公开评论。为保障数据质量,研究团队开展了数据校验流程:剔除空评论条目,并确保数据完整性。随后数据集进入预处理阶段,涵盖文本清洗、采用自然语言工具包(Natural Language Toolkit,NLTK)执行分词、停用词移除,以及借助Sastrawi库进行词干提取。预处理完成后,数据集规模从14742条评论缩减至14051条经校验的评论,以此作为最终用于分析的数据集。本数据集可应用于情感分析、舆情监测、政策评估以及自然语言处理(Natural Language Processing,NLP)研究,能够为探究社交媒体上公众对公共福利项目的反馈提供实证视角。此外,该数据集可重复应用于后续研究、对比分析以及数据驱动型政策评估模型的开发,为公众情感趋势的可持续长期分析提供支撑。

创建时间:
2026-02-23
二维码
社区交流群
二维码
科研交流群
商业服务