遇见数据集

Trace — Cross-Laboratory LLM Steerability Evaluation: responses, judgments, and interpretability data

收藏
Zenodo2026-09-24 更新2026-10-01 收录
官方服务:

资源简介:

Complete results of a cross-laboratory behavioral steerability evaluation of 6 frontier LLMs (24,480 blind peer judgments, leave-one-out consensus), with a mechanistic interpretability extension on Llama-3.3-70B (linear probing and causal activation steering). Raw model responses, all judgment files, steering-sweep generations and judgments, and harvested activations. Version 2 adds the directional ablation experiment on both complementary splits, the reasoning-trace judging for DeepSeek-R1, and the ridge-direction comparison. Cross-model comparisons now use McNemar's exact test rather than Fisher's, reflecting the paired design. Includes all raw per-judge judgment files and the harvested activation tensor.Code, evaluation items, and analysis pipeline: https://github.com/alijalalkamali/trace

提供机构:
Zenodo
创建时间:
2026-09-24
二维码
社区交流群
二维码
科研交流群
商业服务