遇见数据集

The Brain Instruction Set: A Model-Invariant Basis of Semantic Primitives in Large Language Models, from Behaviour to Mechanism

收藏
Zenodo2026-06-15 更新2026-06-18 收录
官方服务:

资源简介:

We extract a small, named, model-invariant set of semantic primitives from large language models, and show they are not merely behavioural regularities but genuine axes of the models' internal representation (mechanistic interpretability; representation universality). Computational test of the Brain Instruction Set (BIS) hypothesis: that human conceptual knowledge can be represented over a finite, shared, internally structured basis of semantic primitives recoverable from large language models. This release contains the BIS Release bundle (152 verified terminals across nine modalities with per-terminal test scores and provenance; schema bis-schema/1.0), the decomposition graph (13,534 concepts / 44,682 weighted edges, GraphML), the cross-model comparison (16-terminal model-invariant core confirmed across an OpenAI and an Anthropic model), the sub-qualia dataset (3,684 stable / 6,928 extended named sub-types with per-terminal dimensionality), the saturation curve, the preprint (SK + EN), figures, and code. v1.1 adds an activation-level result (Phase 15). Probing an open-weight model (Gemma-2-2B) shows the 16-terminal cross-model core is positive on both the correlational and the causal test: variance explained by the BIS basis peaks mid-network (layer 14, R-squared = 0.39 vs 0.007 for a random basis of equal size; inverted-U layer profile), terminals are linearly readable at AUC 0.91-0.95, and norm-calibrated steering of a terminal direction produced a positive projection shift and a changed generation for all four probed terminals. This promotes BIS from a purely behavioural construct toward a candidate mechanistic basis. v1.2 adds the full mechanistic program: scaling to all 152 terminals (R-squared 0.59 at layer 14); cross-architecture replication across Gemma-2-2B, Qwen2.5-1.5B and Phi-3.5-mini (AUC 0.86-0.93 and causal steering replicate; reconstruction R-squared near-saturates for Qwen/Phi due to outlier dimensions); and alignment with independent Gemma Scope sparse-autoencoder features (all 16 core and all 152 terminals match a feature above null; mean cosine 0.44-0.45 vs null 0.09). BIS names roughly 0.3-0.6 percent of the 16,384-feature SAE dictionary -- the shared semantic core that the SAE discovers but cannot itself label or organize. Theoretical outlook (hypothesis, for separate development): that recognizable combinations of BIS terminals may specify the trigger conditions under which innate responses are released -- behavioural, homeostatic, and hormonal -- recasting BIS as a candidate trigger-condition language linking classical ethology (sign stimuli and the innate releasing mechanism) and affective neuroscience. This is an interpretive frame and future work, not a result of the present release. Mandatory declarations: results characterize the structure of textual representations, not phenomenal experience (qualia), not perceptual discriminability, and without external anchors not calibrated truth. Activation-level results are demonstrated across three architectures at small scale; larger-scale replication and a formal steering benchmark are ongoing work.

提供机构:
Zenodo
创建时间:
2026-06-14
二维码
社区交流群
二维码
科研交流群
商业服务