遇见数据集

Climatelaw Turkey: Tweet Corpus and Islamic Terminology Datasets on the 2022 Turkish Climate Law Debate

收藏
Zenodo2026-04-03 更新2026-05-26 收录
官方服务:

资源简介:

This record contains four datasets generated in support of the following publication: Gürlesin, Ömer. "Climate, Religion, and Resistance: Islamic Discourses on Environmental Legislation in Turkish Social Media." [Will be updated]. The datasets document the Islamic discursive framing of the 2022 Turkish Climate Law (İklim Kanunu, Law No. 7339) in Turkish-language social media, specifically on X (formerly Twitter), between January 2023 and May 2025. The corpus (n = 53,438 tweets) was collected using Zeeschuimer (Digital Methods Initiative, University of Amsterdam), a browser-mediated passive capture tool that records publicly visible content rendered in the authenticated user's web browser. This method was adopted following the discontinuation of Twitter's Academic Research API in February 2023 and the introduction of prohibitive commercial pricing (Blakey 2024; Freelon et al. 2024). Collection targeted Turkish-language hashtags and keyword feeds associated with the legislative debate. Data were gathered from publicly accessible accounts only. No user identifiers are retained. Data Collection Tweets were collected using Zeeschuimer (v1.x; Digital Methods Initiative, University of Amsterdam; DOI: 10.5281/zenodo.7525702), a browser extension that captures social media data as it is rendered in the authenticated user's web browser. Unlike automated scrapers, Zeeschuimer performs passive, browser-mediated collection of content the user is already authorized to view. Collection was conducted between January 2023 and May 2025, targeting Turkish-language search results and hashtag feeds on X (formerly Twitter) related to the İklim Kanunu (Turkish Climate Law, Law No. 7339). This methodology was necessitated by the structural inaccessibility of Twitter/X data for academic research following the discontinuation of the Academic Research API in February 2023 and the introduction of enterprise-tier pricing (starting at USD 42,000/month), which rendered platform-official access prohibitive for publicly-funded researchers (Blakey 2024; Freelon et al. 2024; Brown et al. 2025). Browser-mediated collection has since been established as an accepted practice in the post-API research landscape (Rogers 2025). The corpus contains 53,438 unique tweets after deduplication and language filtering (language code = 'tr'). Retweets, original tweets, replies, and quote tweets are retained and labeled in the tweet_type column. Ethical ClearanceThis research was conducted under ethical clearance granted by the Ethics Review Board of the Tilburg School of Catholic Theology (ERB-TST), Tilburg University (The Netherlands), for the research project "Apocalypse and Climate Change" (clearance issued 20 January 2023, valid until project completion). Data collection and processing comply with Tilburg University ethics criteria and applicable Dutch regulations, including the General Data Protection Regulation (GDPR). Anonymization All user identifiers — including usernames, display names, and user IDs — were removed from the dataset prior to archiving. Post IDs (post_id column) are retained, enabling rehydration of tweet records via the X API for researchers with appropriate access credentials. Researchers are advised that tweet text is inherently searchable and that full anonymization of social media content cannot be guaranteed even when direct identifiers are removed. Users are encouraged to follow AoIR and institutional ethics board guidelines when using verbatim tweet text in publications. Sentiment Analysis Turkish sentiment classification was conducted using the savasy/bert-base-turkish-sentiment-cased model (Yildirim 2024), a BERT-based transformer fine-tuned on Turkish product and social media reviews. The model was applied in a zero-shot transfer setting to the climate policy discourse corpus. Labels were mapped to three classes: positive, negative, and notr (neutral). The model achieves F1 = 95.0 on its in-domain benchmark (arXiv:2401.17396). Classification results are discussed in the article body; raw model outputs are not included in this archive. Islamic Terminology Taxonomy A pre-research taxonomy of 464 Islamic and sacralized political terms was compiled from classical Islamic lexica, Quranic vocabulary lists, and prior scholarship on political Islam in Turkey (Yavuz 2003; White 2002; Kuru 2009). Terms were organized into six semantic categories: Theological Core (TH, n=108), Ritual and Devotional Practice (RP, n=87), Moral-Ethical (ME, n=62), Eschatological (ES, n=75), Civil Religion (CR, n=67), and Community and Identity (CI, n=65). The taxonomy was used to guide keyword searches and qualitative reading; 62 terms were ultimately attested in the corpus and are marked in the frequency column of the taxonomy file. Hashtag Coding Top hashtags were identified by raw frequency and grouped by discourse function: Opposition (contesting the law), Campaign (mobilizing signatures or action), General (climate topics without explicit opposition), and Neutral (non-thematic). Orthographic variants (capitalization, Unicode normalization differences) were merged prior to counting. Blakey, Elizabeth. 2024. "The Day Data Transparency Died: How Twitter/X Cut Off Access for Social Research." Journalism Practice 18 (9). https://doi.org/10.1177/15365042241252125 Brown, Megan A., et al. 2025. "Web Scraping for Research: Legal, Ethical, Institutional, and Scientific Considerations." Big Data & Society 12 (1). https://doi.org/10.1177/20539517251381686 Digital Methods Initiative. 2022. Zeeschuimer. University of Amsterdam. https://doi.org/10.5281/zenodo.7525702 Freelon, Deen, Cristina Monzer, Gayoung Jeon, Cameron Moy, and Natasha Williams. 2024. "The Post-API Age of Social Media Data Access: Past, Present, and Future." The Annals of the American Academy of Political and Social Science 716. https://doi.org/10.1177/00027162251372557 Savas Yildirim, “Fine-Tuning Transformer-Based Encoder for Turkish Language Understanding Tasks,” arXiv:2401.17396, preprint, arXiv, January 30, 2024, https://doi.org/10.48550/arXiv.2401.17396.

提供机构:
Zenodo
创建时间:
2026-04-03
二维码
社区交流群
二维码
科研交流群
商业服务