遇见数据集

Dataset — Make Reddit Great Again: Assessing Community Effects of Moderation Interventions on r/The_Donald

收藏
Zenodo2023-01-10 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Reddit contents and complementary data regarding the r/The_Donald community and its main moderation interventions, used for the corresponding article indicated in the title. An accompanying R notebook can be found in: https://github.com/amauryt/make_reddit_great_again <strong>If you use this dataset please cite the related article.</strong> The dataset timeframe of the Reddit contents (submissions and comments) spans from 30 weeks before <em>Quarantine</em> (2018-11-28) to 30 weeks after <em>Restriction</em> (2020-09-23). The original Reddit content was collected from the Pushshift monthly data files, transformed, and loaded into two SQLite databases. The first database, <em>the_donald.sqlite</em>, contains all the available content from r/The_Donald created during the dataset timeframe, with the last content being posted several weeks before the timeframe upper limit. It only has two tables: <em>submissions</em> and <em>comments</em>. It should be noted that the IDs of contents are on base 10 (numeric integer), unlike the original base 36 (alphanumeric) used on Reddit and Pushshift. This is for efficient storage and processing. If necessary, many programming languages or libraries can easily convert IDs from one base to another. The second database, <em>core_the_donald.sqlite</em>, contains all the available content from core users of r/The_Donald made platform-wise (i.e., within and without the subreddit) during the dataset timeframe. Core users are defined as those who authored either a submission or a comment a week in r/The_Donald during the 30 weeks prior to the subreddit's Quarantine. The database has four tables: <em>submissions</em>, <em>comments</em>, <em>subreddits</em>, and <em>perspective_scores</em>. The <em>subreddits</em> table contains the names of the subreddits to which submissions and comments were made (their IDs are also on base 10). The <em>perspective_scores</em> table contains comment toxicity scores. The Perspective API was used to score comments based on the attributes <em>toxicity</em> and <em>severe_toxicity</em>. It should be noted that not all of the comments in <em>core_the_donald</em> have a score because the comment body was blank or because the Perspective API returned a request error (after three tries). However, the percentage of missing scores is minuscule. A third file, <em>mbfc_scores.csv</em>, contains the bias and factual reporting accuracy collected in October 2021 from Media Bias / Fact Check (MBFC). Both attributes are scored on a Likert-like manner. One can associate submissions to MBFC scores by doing a <em>join</em> by the <em>domain</em> column.

提供机构:
Zenodo
创建时间:
2022-02-24
二维码
社区交流群
二维码
科研交流群
商业服务