Replication Data and Code for: Mechanistic Interpretability of Homogeneity in Large Language Models
收藏官方服务:
资源简介:
This repository contains the replication code, Jupyter Notebooks, raw prompt responses, pre-computed embedding manifolds, and layer-by-layer metrics (Participation Ratio and Shannon Entropy) evaluating representations across network depth in frontier Large Language Models (LLMs) compared to human baseline writing. The experiments cover Base vs. Instruct models, SFT vs. DPO post-training paradigms, model scale dynamics, prompt complexity, and Sparse Autoencoder (SAE) feature patching.
提供机构:
Zenodo创建时间:
2026-07-19



