遇见数据集

Religion and Climate Change in the Netherlands (2023): A Multi-Layer Twitter Discourse Dataset Collection

收藏
Zenodo2025-12-12 更新2026-05-26 收录
官方服务:

资源简介:

1. Core Dataset Dataset Name: TwitterClimate2023_CoreData.csvOne-Word Description: Foundation Dataset Description:TwitterClimate2023_CoreData.csv contains the fully cleaned, structured, and anonymized foundation dataset used for all subsequent analyses. The dataset consists exclusively of Dutch-language posts collected from the X platform between 1 January and 31 December 2023. The initial corpus of 770,000 posts was retrieved using Zeeschuimer, a high-capacity scraping tool optimized for large-volume data collection and configured with a targeted keyword protocol centered on the Dutch root klimaat. Following collection, the dataset underwent an extensive Big Data preprocessing pipeline. Procedures included full deduplication, noise removal, text normalization, and a stringent contextual filtering phase designed to exclude posts where klimaat was used in non-environmental senses (e.g., “school climate,” “political climate”). This workflow generated the final analyzable corpus of 698,598 posts, all represented in this dataset. All personal identifiers (usernames, profile names, user IDs) were removed to ensure privacy. The dataset provides a unique item_id for each post, serving as the relational anchor that connects all supplementary datasets in the collection. 2. Sentiment Analysis Dataset Dataset Name: TwitterClimate2023_SentimentScores.csvOne-Word Description: Sentiment Dataset Description:TwitterClimate2023_SentimentScores.csv contains the sentiment classifications assigned to each Dutch-language post in the CoreData dataset. Sentiment analysis was conducted using Python, Pandas, and state-of-the-art RoBERTa-based transformer models (Barbieri et al. 2022) optimized for Dutch social media text. Each post was automatically categorized as positive, neutral, or negative, producing an affective polarity layer required for later triangulation. The sentiment scores provide the quantitative baseline for examining emotional dynamics across grassroots climate discourse.Records link directly to the core dataset via item_id. 3. Religious & Apocalyptic Language Dataset Dataset Name: TwitterClimate2023_ReligiousApocalypticTerms.csvOne-Word Description: ReligiousApocalyptic Dataset Description:TwitterClimate2023_ReligiousApocalypticTerms.csv identifies and categorizes religious, quasi-religious, and apocalyptic terminology appearing across Dutch-language climate-related posts. This analytical layer implements the hierarchical coding framework developed in: Gürlesin, Ömer F. 2025b. Religion and Climate Change in the Netherlands: A Taxonomic Dataset. Zenodo. doi:10.5281/zenodo.14605107. The taxonomy includes macro-categories such as belief systems, apocalyptic narratives, moral–ethical concepts, and symbolic-esoteric expressions. Using this scheme, the dataset captures explicit theological references, implicit moral framings, end-times. All classification rules, codebooks, and operational procedures follow the documentation outlined in Gürlesin (2025c). Each entry is linked to the core dataset via item_id, enabling researchers to correlate lexical–symbolic patterns with sentiment polarity and thematic framing.This dataset forms the interpretive-symbolic dimension of the computational mixed-methods framework. 4. Climate Theme Classification Dataset Dataset Name: TwitterClimate2023_ClimateThemesClassified.csvOne-Word Description: Themes Dataset Description: TwitterClimate2023_ClimateThemesClassified.csv contains the thematic categorization of Dutch-language climate posts based on compound climate-related lexical formations. The thematic identification is structured according to: Gürlesin, Ömer F. 2025a. Dutch Compound Climate-Related Terms Taxonomy (2023). Zenodo. doi:10.5281/zenodo.17724263. This taxonomy, developed as part of the methodological design of the study, enables the systematic grouping of posts into themes such as activism, governance, skepticism, crisis framing, ethics, environmental solutions, and economic dimensions. Compound word groups such as klimaatplicht, klimaatbeleid, and klimaatactie were algorithmically detected and used as thematic markers. The dataset is directly linked to the core file via item_id, enabling cross-layer integration with sentiment polarity and religious/apocalyptic terminology.This classification represents the thematic–structural dimension of the mixed-methods triangulation strategy. Collection Overview This dataset collection—composed of CoreData, SentimentScores, ReligiousApocalypticTerms, and ClimateThemesClassified—implements a Computational Mixed-Methods design applied to Dutch-language climate discourse on Twitter/X in 2023. The methodological framework consists of: Big Data collection via Zeeschuimer, Extensive preprocessing yielding 698,598 clean climate-relevant posts, Transformer-based sentiment modeling (RoBERTa), Two specialized taxonomies: Religious/apocalyptic taxonomy → used in dataset 3 (Gürlesin 2025b) Compound climate-term taxonomy → used in dataset 4 (Gürlesin 2025a), Visualization, co-occurrence, and exploratory analysis in Power BI, Integration of all analytical outputs through a shared item_id architecture.

提供机构:
Zenodo
创建时间:
2025-12-12
二维码
社区交流群
二维码
科研交流群
商业服务