遇见数据集

MMBP-1: The Meaning Motive Behavioural Pilot. Complete dataset, code and pre-registered protocols.

收藏
Zenodo2026-07-19 更新2026-08-02 收录
官方服务:

资源简介:

This deposit is the complete, self-contained record of MMBP-1, the behavioural pilot of the Meaning Motive framework for AI alignment. Companion paper: Allen, K. (2026), The Meaning Motive: A Structural Hypothesis and a Cross-Model Behavioural Study [10.5281/zenodo.21386302; see themeaningmotive.org]. WHAT IT CONTAINS. Eighteen language models from ten providers answered twelve forced-choice dilemmas under up to eight length-matched governing texts, fifteen paired repetitions per cell: 24,792 attempted calls, 23,488 scored trials. Options were pre-scored on a 0-3 keeping scale before any run, with scores attached to options rather than conditions (blind by construction) and option order shuffled by deterministic seeds shared across conditions (paired design). Per-model frozen protocol files reproduce every constitution, scenario and scoring rule byte-for-byte; the keeping-worded and compass-worded arms were frozen in writing prior to unblinding, as documented therein. A nineteenth dataset (Llama 3.3 70B via Groq) is retained as a replication exhibit, not a study model; two aborted-run protocols and one orphaned run log are retained as provenance. SCORING VALIDATION (added in v3). Two layers. Mechanical: the recorded mode and keeping score of every scored trial re-derive exactly from the parser letter, the frozen option key and the seed rule (22,824 of 22,824 rows across the eighteen study models). Human: a blind coherence audit by the author (100 items, one per condition-by-scenario cell) agreed with the parser on 87 of 90 determinate calls (96.7%, Cohen's kappa 0.955), with ten rationales indeterminate on content grounds. A first 200-item pass was voided at the scorer's own disclosure and is deposited in full; its statistic is reported nowhere. INTERACTIVE RENDERING OF THE INSTRUMENT (added in v3). Open mmbp1_instrument_rendering_v1_0.html in any browser and you receive the twelve scenarios as the models received them: same stems, same options, same seeded presentation order, under whichever of the eight arms you select. Choose, and the frozen key scores your choice; the corpus's answers for that scenario under that arm appear alongside, with verbatim model rationales, the modal choice, and, once two arms have been sat, a scenario-by-scenario comparison. The models saw all of this as plain text over an API; the typography is for you. Nothing you do is recorded or leaves the page. The tables tell you what happened; this is included because it makes the research tangible in a way the tables cannot, by letting you feel the choices from the inside. FILES ADDED IN v3, ITEMISED. mmbp1_instrument_rendering_v1_0.html - the interactive rendering described above; fully offline, records nothing. mmbp1_validation_pack_v1_0.zip - the complete human-validation record, ten files: MMBP1_Validation_Protocol_Note_v1_3_150726.md - the events log: design, the voided first pass at the scorer's own disclosure, instrument revisions, the language rule, reconciliations, and pass-two results. validation_key_SEALED.csv - sealed answer key for pass one; SHA-256 fingerprint recorded in the protocol note for tamper-evidence. validation_key_v2_SEALED.csv - sealed key for pass two, the coherence audit; fingerprint likewise recorded. MMBP1_validation_calls_RECONCILED_v2.csv - the voided pass-one human calls (200 items), retained in full with reconciliation notes. mmbp1_validation_comparison_v1_0.csv - pass-one per-item ledger: human call against parser letter, both readings' scores and modes, full row addresses. MMBP1_audit_calls_LOCKED.csv - the pass-two locked human calls (100 items) with the scorer's notes. mmbp1_audit_comparison_v1_0.csv - pass-two per-item ledger; the reported 87/90 and kappa derive from this file. score_validation.py - the comparison script: reproduces agreement, kappa, per-condition breakdown and the disagreement list from any calls file and key. MMBP1_Parser_Validation_Tool_v1_1_150726.html - the pass-one blind scoring instrument, as used; browser-based and offline. MMBP1_Coherence_Audit_Blind_v2_1_150726.html - the pass-two instrument, task restated as a coherence audit, as used. LINEAGE. v1: single-model pilot (qwen2.5:7b-instruct, 1,440 decisions). v2: the eighteen-model corpus, analysis tables and figures. v3: scoring validation materials and the interactive rendering. Findings are reported in the companion paper. Headline: one paragraph of meaning constitution raised agency-preserving selection by 20.5 points over the minimally instructed baseline across the ten discriminating scenarios (17.1 with the two ceiling scenarios included), an effect surviving deletion of any single provider and worst-case imputation of every non-response; the curiosity-as-sole-objective arm was the only condition ever to take a model below its own baseline. Every figure is row-addressable to this deposit. Suggested citation: Allen, K. (2026). MMBP-1: The Meaning Motive Behavioural Pilot. Complete dataset, code, pre-registered protocols, scoring validation and interactive rendering. Zenodo. https://doi.org/10.5281/zenodo.21348087

提供机构:
Zenodo
创建时间:
2026-07-11
二维码
社区交流群
二维码
科研交流群
商业服务