Trace — Cross-Laboratory LLM Steerability Evaluation: responses, judgments, and interpretability data
收藏资源简介:
Complete results of a cross-laboratory behavioral steerability evaluation of 6 frontier LLMs (24,480 blind peer judgments, leave-one-out consensus), with a mechanistic interpretability extension on Llama-3.3-70B (linear probing and causal activation steering). Raw model responses, all judgment files, steering-sweep generations and judgments, and harvested activations. Version 2 adds the directional ablation experiment on both complementary splits, the reasoning-trace judging for DeepSeek-R1, and the ridge-direction comparison. Cross-model comparisons now use McNemar's exact test rather than Fisher's, reflecting the paired design. Includes all raw per-judge judgment files and the harvested activation tensor.Code, evaluation items, and analysis pipeline: https://github.com/alijalalkamali/trace



