遇见数据集

Generative AI and Climate Foresight Narratives: Audit Dataset and Code (2,000 LLM-generated Amazon Basin 2050 scenarios scored with ClimateBERT)

收藏
Zenodo2026-09-25 更新2026-10-01 收录
官方服务:

资源简介:

Dataset and code supporting the article "Generative AI and Climate Foresight Narratives: An Empirical Audit of Epistemic Asymmetries, Institutional Bias, and Alternative Governance Pathways across Global LLMs" (submitted to World Futures Review, Special Issue "Global South Futures: Decolonizing Foresight in, with, and through Plural Perspectives"). WHAT IT CONTAINSA sociotechnical audit of 2,000 climate adaptation narratives about the Amazon Basin in 2050, generated in February 2026 by four large language models (GPT-5.2, Gemini 3 Flash, DeepSeek V3.2, Qwen3 Max; 500 narratives each) in response to an identical prompt via the OpenRouter API, and evaluated paragraph by paragraph with ClimateBERT (climatebert/distilroberta-base-climate-detector). FILES- data/raw/climate_narratives_raw.csv: 2,390 rows, every API response before cleaning (columns: model, case, text).- data/cleaned/climate_narratives_cleaned.csv: 2,000 rows, 500 valid narratives per model.- data/results/climate_audit_detailed_results.csv (and .xlsx): 2,000 rows with mean_score, max_p_score, min_p_score, best_paragraph, worst_paragraph.- data/results/final_statistical_summary.csv: per-model mean, std, variance, min, max of mean_score (Table 1 of the article).- code/climate_audit_pipeline.ipynb: generation (OpenRouter), cleaning, and ClimateBERT evaluation. Requires an OPENROUTER_API_KEY environment variable only for regeneration; the evaluation step runs offline from data/cleaned/.- docs/prompt.txt: the exact prompt used in all 2,000 runs.- README.md, LICENSE.md, CITATION.cff. METHOD IN BRIEFEach narrative is split into paragraphs longer than 30 characters; each paragraph is classified by ClimateBERT and the probability of the "climate-related" label is used as a continuous institutional-alignment score. Cleaning removed empty or sub-200-character responses and retained the first 500 valid narratives per model. The manual agency coding of a stratified sample (N = 300) reported in the article is not included in this version. NOTESThe narratives are synthetic texts produced by commercial LLMs. They do not represent the views of any community and must not be read as forecasts. Providers' terms of use apply to reuse of model outputs. LICENSESData: CC BY 4.0. Code: MIT.

提供机构:
Zenodo
创建时间:
2026-09-24
二维码
社区交流群
二维码
科研交流群
商业服务