遇见数据集

YouTube Discourse on Indonesia’s Makan Bergizi Gratis Policy

收藏
Mendeley Data2026-08-08 收录
官方服务:

资源简介:

This dataset supports the study “Normative Support and Fragile Legitimacy: A Computational Analysis of Public Discourse on Indonesia’s Makan Bergizi Gratis Policy.” It contains YouTube video metadata, publicly visible comment data, and the Python script used to collect the source material for computational analysis of digital policy discourse. The repository includes: (1) Youtube_Videos_makan_bergizi_gratis_FULL.csv, containing 276 video records from 77 channels; (2) Youtube_Comments_makan_bergizi_gratis_FINAL.csv, containing 35,450 comment records associated with 255 source-video titles; and (3) youtube_scraper.py, which implements keyword-based video discovery, metadata retrieval, channel-size filtering, comment extraction, and CSV/Excel export using yt-dlp and pandas. The scraper was configured with the search phrase “makan bergizi gratis,” a requested maximum of 1,000 search results, a minimum channel size of 10,000 subscribers, and a video-publication-date filter from 1 January 2024 to 31 December 2025. The archived video records range from 24 May 2024 to 31 December 2025, while stored comment dates range from 26 January 2025 to 26 January 2026 because the script filters video publication dates rather than comment-posting dates. Video-level variables include title, URL, channel name, subscriber count, publication date, view count, and platform-reported comment count. Comment-level variables include source-video title, original comment text, displayed username, comment date, like count, detected reply count, and parent-comment information. In the deposited output, all comments are marked as main posts and all reply counts are zero; therefore, this version should not be used for reply-network analysis. The data represent a keyword-based, platform-mediated, non-probability sample and must not be interpreted as representative of Indonesian public opinion. Search ranking, channel eligibility, user self-selection, moderation, deleted content, and comment ordering may affect the observed discourse. Engagement statistics are snapshots captured at collection time and may subsequently change. The comments are predominantly Indonesian and may include informal spelling, abbreviations, emojis, regional expressions, code-switching, sarcasm, and political language. The dataset is intended for research on computational social science, public-policy communication, policy legitimacy, digital public opinion, topic modelling, sentiment analysis, and Indonesian-language social-media discourse. Because the raw data contain usernames and searchable verbatim comments, users must not attempt to identify, contact, profile, or target commenters. Public release and reuse should follow applicable ethical, legal, institutional, privacy, and platform requirements. *) For more details, please check the README.md file.

创建时间:
2026-07-29
二维码
社区交流群
二维码
科研交流群
商业服务