lbe — Cross-Laboratory LLM Steerability Evaluation: responses, judgments, and interpretability data
收藏官方服务:
资源简介:
Complete results of a cross-laboratory behavioral steerability evaluation of 6 frontier LLMs (24,480 blind peer judgments, leave-one-out consensus), with a mechanistic interpretability extension on Llama-3.3-70B (linear probing and causal activation steering). Raw model responses, all judgment files, steering-sweep generations and judgments, and harvested activations. Code, evaluation items, and analysis pipeline: https://github.com/alijalalkamali/lbe
提供机构:
Zenodo创建时间:
2026-07-27



