遇见数据集

Can Psychometric Trait Vectors Be Injected into Neural Network Weights? A Controlled Empirical Test and a Falsifiable Roadmap

收藏
Zenodo2026-08-13 更新2026-08-20 收录
官方服务:

资源简介:

A recurring proposal in the human-AI alignment literature is that quantified human psychological traits can be "injected" into a trained model's parameters by projecting a low-dimensional trait vector onto the weight space of a target layer. We formalize this proposal, which we term trait-vector weight injection, as two explicit, falsifiable hypotheses, and subject it to a controlled empirical test that proposals of this kind typically omit: a magnitude-matched random-vector control condition. Across N=40 independent training runs on a supervised image-classification task, a fixed trait vector (creativity, empathy, metacognition; nominally scaled to the normal range of standard psychometric instruments) projected through a random matrix onto a hidden-layer weight matrix produces a small but statistically significant decrease in test accuracy relative to an untouched baseline (delta = -0.24 percentage points, paired t(39) = -3.37, p = 0.0017, Cohen's d = -0.53). Critically, this effect is statistically indistinguishable from the effect of a magnitude-matched random vector projected through the same matrix (delta = -0.10 points, t(39) = -1.39, p = 0.17; Wilcoxon p = 0.46; bootstrap 95% CI for the difference [-0.24, +0.03] points). A five-point sensitivity sweep over the injection magnitude eta in {0.01, 0.02, 0.05, 0.10, 0.20} (N=20 independently seeded runs per level) confirms that the cognitive-specific and random-control conditions remain statistically indistinguishable at every tested magnitude, while the sign and significance of the cognitive-vs-baseline comparison itself is unstable across independent resamples (the eta = 0.05 sweep estimate, drawn from a separate set of 20 seeds, is not a replication of the N=40 main experiment at the same nominal magnitude, and the two should not be compared as if they were). A multi-direction robustness check, comparing the cognitive vector against five independent random directions per seed rather than one, corroborates the null result: the cognitive vector's effect is statistically indistinguishable from a typical draw from the random-direction distribution (p = 0.15; 58.5% of random directions are at least as extreme). We conclude that, in this form, trait-vector weight injection via untrained random projection of psychometric test averages carries no detectable information content beyond that of generic parameter noise of matched magnitude, and we present a falsifiable, testable roadmap, grounded in the validated model-editing and activation-steering literature, for what a genuine test of human-cognitive weight infusion would require. We regard this as a legitimate contribution in its own right: a pre-registered-style negative result with an explicit control condition, clarifying which formulations of "symbiotic" human-AI integration are and are not empirically supported.

提供机构:
Zenodo
创建时间:
2026-08-13
二维码
社区交流群
二维码
科研交流群
商业服务