遇见数据集

CCF Database: A Machine-Learning-Annotated Corpus of 283,964 Canadian Climate Articles (1978–2026) — PostgreSQL edition

收藏
Zenodo2026-08-14 更新2026-08-20 收录
官方服务:

资源简介:

The Canadian Climate Framing (CCF) Database v2.0.0: a machine-learning-annotated, sentence-level corpus of 283,964 Canadian climate-change articles from 22 news outlets (1978–2026), processed into 9,908,776 two-sentence analytical units, each carrying 65 hierarchical binary annotations, named-entity extractions, per-article aggregates, per-category reliability tiers, and 10,192,740 BAAI/bge-m3 sentence embeddings. Release notes (v2.0.0): corpus extended to 31 July 2026 under the stewardship of the CCF Observatory (continuous extraction–annotation pipeline; the database is now refreshed on Zenodo approximately every six months); two news websites of the Canadian public broadcaster (CBC.ca, Radio-Canada.ca, covered from 2023) join the corpus; two English security sub-category models were restored after a storage migration and all affected recent rows were re-annotated; an English tokenizer-casing defect in the live annotation service was found and fixed, and every English annotation of the post-2024 continuous corpus was re-computed with the corrected pipeline. The validation layer (training data, models, gold standard, reliability tiers) is unchanged from v1.1.0, and all v1.1.0 rows carry over bit-identically. The full methodology is described in the accompanying paper (Scientific Data, in revision). Because Canadian newspaper articles are protected by copyright, the deposit does not redistribute full article text; every row is traceable to its source article through complete bibliographic metadata.

提供机构:
Zenodo
创建时间:
2026-08-14
二维码
社区交流群
二维码
科研交流群
商业服务