遇见数据集

Climate Change Media Coverage Dataset (2011–2026)

收藏
Zenodo2026-06-26 更新2026-06-17 收录
官方服务:

资源简介:

Climate Change Media Coverage Dataset (2011–2026) Description This dataset contains a structured corpus of 49,627 newspaper articles examining media coverage of climate-related issues between 2011 and 2026. The corpus was designed to support comparative media analysis across countries and news outlets with different political orientations, with a particular focus on discourse, framing, and representations of climate change. The dataset covers several major thematic dimensions, including: Climate change and climate crises International climate summits (COP conferences) Climate activism and protest movements Environmental policy and institutions Energy transitions, emissions, and sustainability The corpus includes articles from mainstream liberal/centrist and conservative media outlets in the following countries: France Germany Greece Hungary Italy Poland Spain Sweden The data covers a time span from 2011 to 2026. Data Collection Articles were collected through automated web scraping using language-specific keyword queries related to climate change and environmental issues. Keywords were adapted to each language. Search queries were standardized across countries to maximize comparability. Retrieved articles were filtered and cleaned to remove duplicates, noise, and irrelevant entries. To facilitate comparative analysis, data collection focused on three broad thematic domains: International climate summits (e.g., COP conferences) Climate activism and protest movements Environmental policy developments and institutions Data Structure The dataset is provided in tabular format and includes the following variables: Article_URL — URL of the original article Title — Original article title Title_eng — Article title translated into English Date — Publication date Year — Publication year Section_eng — Newspaper section translated into English Site — Media outlet Country — Country of publication Orientation — Political orientation of the media outlet (e.g., liberal, conservative) Language — Language of publication Keyterms — Climate-related keywords identified for the article Due to copyright restrictions, the dataset does not include full article texts. Preprocessing The following preprocessing steps were applied: Removal of duplicates and non-relevant entries Cleaning of HTML artifacts and boilerplate content Normalization of metadata fields Translation of selected metadata fields into English Harmonization of keyword variants across languages Limitations Several limitations should be considered when using the dataset: The corpus reflects the editorial practices and agendas of selected media outlets and is not fully representative of national media systems. Automated translation may introduce minor semantic distortions. Keyword-based collection may omit relevant articles that do not contain the selected search terms. Differences in archive availability and publication practices across media outlets may affect coverage. Ethical Considerations This dataset is intended for academic and research purposes. Users are responsible for complying with applicable copyright regulations and the terms of use of the original news sources. No personal or sensitive data was intentionally collected.

提供机构:
Zenodo
创建时间:
2026-06-16
二维码
社区交流群
二维码
科研交流群
商业服务