遇见数据集

Source data — Calibrating LLM-Derived Trust Scores for News Outlets (Pilot A, 52 outlets)

收藏
Zenodo2026-06-22 更新2026-06-28 收录
官方服务:

资源简介:

Frozen source data for the pilot study Calibrating LLM-Derived Trust Scores for News Outlets When Public Factuality Scorecards Disappear (Claassen & van Vuuren). Fifty-two English-language news outlets were scored with a fixed LLM pipeline and compared against a frozen Media Bias Fact Check (MBFC) factuality snapshot (June 2026). MBFC labels were used only for offline calibration and evaluation, never in prompts or crawls. This deposit includes: outlet-level scores and GAP metrics for all pipeline steps; train/validation split (42/10, seed 42); MBFC reference labels and provenance; affine calibration parameters (Steps 6–7); crawl quality audit; Wikipedia/search provenance tables; NELA-GT-2020 secondary triangulation (37-outlet overlap); frozen per-outlet LLM and web-crawl snapshots; data dictionary; and a Reuters worked example (docs/reuters-golden-thread.md). Start with README.md and DATA-DICTIONARY.md inside the archive.

提供机构:
Zenodo
创建时间:
2026-06-22
二维码
社区交流群
二维码
科研交流群
商业服务