Securitizing Climate Change: Processed Corpora of U.S. Presidential Speeches, Congressional Research Service Reports, and Legislative Bills, 2003–2023
收藏资源简介:
This dataset contains eight files supporting the analysis in "The War on Climate Change: A Computational Analysis of Presidential and Congressional Climate Security Discourse, 2003–2023." The files comprise three climate-related document corpora and their securitizing subsets: (1) U.S. presidential speeches collected from the American Presidency Project across eleven speech genres; (2) Congressional Research Service report summaries; and (3) congressional bill summaries (HR, HJRes, S, SJRes), with the latter two sourced via the Congress.gov API. Climate-related documents were identified through a two-stage filtering process combining a keyword-based climate lexicon and the ClimateBERT transformer model (huggingface.co/kruthof/climateattention-10k-upscaled). Securitizing documents were classified using a security lexicon adapted from Baele and Sterck (2014), retaining documents whose security ratio exceeded a corpus-specific baseline. Each record includes the document text, metadata, and computed security ratio. The dataset also includes the climate change lexicon used for initial document filtering and the security lexicon used to calculate document security ratios, provided as standalone files to support replication.



