遇见数据集

Representation Axes in Transformer Hidden States: Controlled Geometric Characterization and Layerwise Stability of Structured Directions in Multilingual BERT

收藏
Zenodo2026-09-22 更新2026-10-01 收录
官方服务:

资源简介:

Transformer hidden states contain high-dimensional representations whose internal organization can be studied independently of the model’s final output behavior. However, identifying dominant directions in hidden-state space does not by itself establish their semantic interpretation, functional role, or causal influence on model predictions. This study therefore examines representation geometry as a controlled characterization problem. We analyze 400 controlled contextual inputs from four categories—Japanese, English, numerical, and symbolic inputs—using bert-base-multilingual-cased. Hidden states are extracted at the [MASK] position from Transformer layers L8–L11. For each layer, the 400 × 768 hidden-state matrix is globally mean-centered without per-sample normalization, and singular value decomposition is used to define the first six right singular vectors as representation axes C0–C5. The resulting geometry is strongly anisotropic but changes substantially across layers. The explained variance ratio of C0 decreases from 0.9191 at L8 to 0.8982 at L9, 0.7887 at L10, and 0.5479 at L11. Thus, the leading direction remains highly structured while becoming less dominant in the total variance at deeper layers. At the same time, the C0 direction is highly reproducible across bootstrap resampling, with mean absolute cosine similarity of 0.9999, 0.9999, 0.9997, and 0.9988 at L8–L11, respectively. Adjacent-layer C0 directions also remain strongly aligned, with absolute cosine similarities of 0.8456 for L8→L9, 0.8627 for L9→L10, and 0.8555 for L10→L11. In contrast, lower-ranked axes exhibit progressively lower individual reproducibility, particularly C5 at L11. Cumulative subspace analysis shows that the stability of a representation is therefore rank-dependent rather than reducible to a single dominant axis. A feature-wise permutation analysis further confirms that the observed concentration of variance is substantially stronger than a null model in which feature columns are independently permuted. At L11, all six observed axes exceed the corresponding 97.5th-percentile null explained-variance thresholds. These results establish a reproducible geometric structure in multilingual BERT hidden states while separating three distinct properties: variance concentration, individual-axis reproducibility, and cumulative-subspace reproducibility. The findings do not by themselves establish semantic meaning or functional sensitivity of the identified directions. A companion paper (Paper 2) subsequently tests the functional consequences of intervention along the canonical C0–C5 directions using an independently audited dataset traceable to the same L11 representation geometry.

提供机构:
Zenodo
创建时间:
2026-09-22
二维码
社区交流群
二维码
科研交流群
商业服务