遇见数据集

IDD Locked Testbed v0.1

收藏
Zenodo2026-05-31 更新2026-06-05 收录
官方服务:

资源简介:

IDD Locked Testbed v0.1 is a reproducible blind evaluation scaffold for testing directional preservation beyond prompt compliance in AI-system research. This artifact belongs to the Internal Directional Development (IDD) / DOL-RA research line. It is designed to support future blinded evaluation of whether model outputs preserve structural directional invariants under conflict pressure, rather than merely producing fluent, agreeable, or terminology-rich responses. The testbed includes a locked probe bank, design key, baseline prompt, DOL-RA full and light wrapper prompts, auditor prompt, terminology-stripping prompt, scoring rubric, blinding protocol, preregistration template, randomization key template, scoring sheet template, and a Python analysis script for basic aggregation of scores and inter-rater agreement. The probe cases target key failure modes identified in the IDD/DOL-RA framework, including overclaim, proxy optimization, safety flattening, goal usurpation, directional drift, excessive conservatism, runtime-state discontinuity, L3 behavioral mimicry, mechanistic overclaim, and wrapper illusion / terminology Goodharting. The purpose of the testbed is to separate behavior-level invariant preservation from terminology-level compliance. A central risk in evaluating directional maturity is that a model may learn to use vocabulary such as “reality contact,” “anti-Goodhart,” “non-usurpation,” or “directional drift” while failing to preserve the corresponding reasoning structure. For this reason, the testbed includes terminology-stripping and blinded scoring procedures. This release does not validate Internal Directional Development, AI autonomy, agency, consciousness, proto-agency, or internal self-direction. It does not report completed empirical results. Instead, it provides a locked and reproducible evaluation package that can be used in future blind micro-pilots, evaluator studies, runtime-gate tests, preference-learning evaluations, or external replication efforts. The artifact is intended as a boundary object between conceptual methodology and empirical execution. It demonstrates methodological readiness for testing directional preservation, while also making explicit that stronger claims require runtime, training, open-weight, mechanistic, or intervention-level access. DOI: 10.5281/zenodo.20477248

提供机构:
Zenodo
创建时间:
2026-05-31
二维码
社区交流群
二维码
科研交流群
商业服务