CCF Database: A Machine-Learning-Annotated Corpus of 283,964 Canadian Climate Articles (1978–2026) — PostgreSQL edition
收藏资源简介:
The Canadian Climate Framing (CCF) Database: a machine-learning-annotated, sentence-level corpus of 283,964 Canadian climate-change articles from 22 news outlets (1978–2026), processed into 9,908,776 two-sentence analytical units, each carrying 65 hierarchical binary annotations, named-entity extractions, per-article aggregates, per-category reliability tiers, and 10,192,740 BAAI/bge-m3 sentence embeddings. Release notes (v2.0.1 — documentation and lookup-table fixes; the corpus itself is unchanged from v2.0.0): (i) the exclusion_reason column of CCF_reliability_tiers now stores true NULLs instead of the literal string “nan” on the 63 rows without an exclusion; (ii) the archived static training configuration was re-derived from the released training code, correcting four parameter descriptions (reinforcement trigger F1 < 0.70; reinforced-phase learning rate 5e-6; best-epoch criterion 0.98·F1₁ + 0.02·macro-F1; unweighted cross-entropy in the normal phase, WeightedRandomSampler + class-weighted cross-entropy [1.0, 3.5] in the reinforced phase); (iii) the English base model is correctly identified as bert-base-cased (tokenizer must be loaded with do_lower_case=False); (iv) the attached methodology PDFs and code bundle reflect these corrections, and the LaTeX sources in the bundle now ship corpus_macros.tex so they compile as archived. All row counts, annotations, embeddings and validation metrics are identical to v2.0.0. The full methodology is described in the accompanying paper (Scientific Data, in revision). Because Canadian newspaper articles are protected by copyright, the deposit does not redistribute full article text; every row is traceable to its source article through complete bibliographic metadata.



