遇见数据集

HAI / BIS: A Finite, Shared, Structured Basis of Semantic Primitives Extractable from Large Language Models

收藏
Zenodo2026-06-14 更新2026-06-17 收录
官方服务:

资源简介:

Computational test of the Brain Instruction Set (BIS) hypothesis: that humanconceptual knowledge can be represented over a finite, shared, internallystructured basis of semantic primitives recoverable from large language models.This release contains the BIS Release bundle (152 verified terminals across ninemodalities with per-terminal test scores and provenance; schema bis-schema/1.0),the decomposition graph (13,534 concepts / 44,682 weighted edges, GraphML), thecross-model comparison (16-terminal model-invariant core confirmed across anOpenAI and an Anthropic model), the sub-qualia dataset (3,684 stable / 6,928extended named sub-types with per-terminal dimensionality), the saturation curve,the preprint (SK + EN), figures, and code. v1.1 adds an activation-level result (Phase 15). Probing an open-weight model(Gemma-2-2B) shows the 16-terminal cross-model core is positive on both thecorrelational and the causal test: variance explained by the BIS basis peaksmid-network (layer 14, R-squared = 0.39 vs 0.007 for a random basis of equalsize; inverted-U layer profile), terminals are linearly readable at AUC0.91-0.95, and norm-calibrated steering of a terminal direction produced apositive projection shift and a changed generation for all four probedterminals. This promotes BIS from a purely behavioural construct toward acandidate mechanistic basis. Mandatory declarations: results characterize the structure of textualrepresentations, not phenomenal experience (qualia), not perceptualdiscriminability, and without external anchors not calibrated truth. Theactivation-level result is a single model at one scale; breadth acrossarchitectures is ongoing work.

提供机构:
Zenodo
创建时间:
2026-06-14
二维码
社区交流群
二维码
科研交流群
商业服务