遇见数据集

Reddit and BlueSky Observations pertaining to Antisemitism Research (June 2024 to August 2025)

收藏
Zenodo2025-12-03 更新2026-05-26 收录
官方服务:

资源简介:

This dataset accompanies the Master of Statistics thesis “Temporal Profiling of Antisemitism on Social Media (June 2024 – August 2025)” and contains the full set of Reddit and BlueSky text observations used for large-scale longitudinal analysis. The data were collected, processed, and annotated between June 2024 and August 2025 using custom pipelines built with AWS Lambda, Step Functions, and platform-specific extraction protocols. All data were cleaned and prepared exclusively for academic research into antisemitic discourse, temporal variation, and platform-specific dynamics. The associated thesis and manuscript are forthcoming. A link to the published version (arXiv) will be added to this record upon release. The upload includes three files: 1. labelled_reddit_sample_observations.csv A manually curated and model-assisted labelled dataset containing 2,413 Reddit observations. These samples were generated through a stratified sampling regimen and annotated using a hybrid workflow combining GPT-5 and Perplexity Pro classification, followed by iterative human verification.Columns include: observation_month – Month of the Reddit observation subreddit – Source subreddit theme – Assigned thematic bucket (e.g., overt, conspiracy, identity, zionism) keyword – Extracted keyword used to map the observation cleaned_text – Pre-processed text used for classifier development label – Binary indicator of antisemitic content (1 = antisemitic, 0 = non-antisemitic) This file served as the core training and audit sample for iterative classifier refinement (HateBERT domain adaptation, Audit 1–3). 2. reddit_observations_jun24_to_aug25.pkl A comprehensive dataset of 171,000 Reddit observations, collected using a three-dimensional stratification schema (month × subreddit group × keyword theme) designed to ensure balanced and representative coverage across time periods, communities, and thematic categories.The dataset includes posts, comments, and replies collected from June 2024 to August 2025.Columns include (abridged): observation_month, subreddit, subreddit_group, theme, keyword Timestamps (created_utc) Metadata ( post_title, scores for posts/comments/replies) Full text fields (body_text, comment_text, reply_text, combined text) Derived features (word_count, cleaned_text) This file underpins the Reddit-based temporal trend analysis and classifier inference presented in the thesis. 3. bluesky_posts_jun24_to_aug25.pkl A complete corpus of 2,717,306 BlueSky posts, collected via a two-layer network expansion protocol seeded from a curated list of journalist, academic, community, progressive, conservative, and far-right accounts.The dataset captures all public text posts authored or reshared by accounts reachable within two network hops of the seed accounts between June 2024 and August 2025.Columns include: created_at, text, Pre-processed text (cleaned_text) Keyword-based thematic tags (keyword_tags) Boolean thematic indicators (overt_flag, conspiracy_flag, zionism_flag, identity_flag) This dataset forms the basis for the BlueSky temporal antisemitism analysis, thematic prevalence estimates, and keyword-mapped trend comparisons. Usage and Licensing All data have been processed to remove platform-specific identifiers where possible, retaining only information necessary for reproducible academic study. The dataset is intended solely for research, auditing, and methodological replication. Users should comply with the respective platform terms of service and relevant ethical guidelines when using these data.

提供机构:
Zenodo
创建时间:
2025-12-03
二维码
社区交流群
二维码
科研交流群
商业服务