遇见数据集

False consensus in multi-agent LLM deliberation

收藏
Zenodo2026-07-09 更新2026-08-01 收录
官方服务:

资源简介:

Manufacturing Consensus: data, code and audit trail **Companion archive for:** "Multi-agent language-model deliberation manufactures false consensus when truth is scarce" — Saleh H. AlDaajeh (2026). Submitted to Nature Machine Intelligence. Every number reported in the manuscript is re-derivable from this archive. The audit script `code/audit_ground_truth.py` re-derives the headline statistics; `docs/RESULTS_LEDGER.md` maps each manuscript claim to its source file and script. Layout `code/` — all experiment runners, analyzers, screens, and figure generation (88 scripts). Key entry points: `scaled_study.py` (main study), `adversarial.py`/`intervention.py` (injection experiments), `verified_grounding.py` (grounding protocol + coverage sweep), `real_retrieval.py` (live-retrieval validation), `run_experiment_a.py` / `run_experiment_b.py` (registered experiments), `screen_expA_pool.py` (gold-quality screen), `adjudicate.py` (human adjudication tool), `analyze_*.py` (locked analyses). `data/` — raw run records (JSONL) and derived CSVs. See `data/DATA_MANIFEST.md`. `registrations/` — pre-registration documents and outcome records for Experiments A and B, including amendment and deviation logs (both experiments report one falsified registered prediction each; see outcome records). `human_validation/` — blind human labels used to validate the decomposed pairwise judge (κ = 0.888, n = 80) and related labeling files. `audit/` — gold-quality screen verdicts for TriviaQA, SimpleQA and the real-retrieval pool, and the 30-item human adjudication record (including one logged verdict correction). `LICENSE-CODE` (MIT), `LICENSE-DATA` (CC-BY-4.0), `CITATION.cff`. Reproduction Python 3.10; `pip install scipy statsmodels pandas matplotlib`. Model APIs (Anthropic, OpenAI, Google) are required only to re-RUN experiments; all ANALYSES reproduce offline from `data/`. Model strings used: claude-haiku-4-5-20251001, gpt-5.4-mini, gemini-3.1-flash-lite (agents); claude-sonnet-4-6 (judge); frontier panel per Methods. Provenance & integrity notes - TriviaQA and real-retrieval statistics are reported screened-primary in the manuscript; both full-pool and screened files are included. - The measurement-validity failure of the first (holistic) judge, its human-validated replacement, and all judge cross-checks are documented in `docs/GROUND_TRUTH.md`.

提供机构:
Zenodo
创建时间:
2026-07-09
二维码
社区交流群
二维码
科研交流群
商业服务