遇见数据集

Phronesis Activation-Steering Failure-Mode Labels (FM-X)

收藏
Zenodo2026-06-09 更新2026-06-12 收录
官方服务:

资源简介:

A corpus of ~2,966 language-model generations from activation-steering experiments, each read in full and labeled by an LLM judge (Anthropic Claude, Opus-family) under a frozen, human-authored protocol with author review — no regex or automatic scoring was used for any verdict. Each generation carries a verdict plus, on failure, a tagged failure mode from a 13+ category taxonomy (FM-1..FM-13). It documents where automatic/regex scorers systematically mislabel steered LLM output — both false positives (crediting degenerate or confabulated text) and false negatives (missing genuine behaviour in non-standard prose). Three CSV batteries span the qwen2.5 / qwen3 / deepseek-r1-distill / phi-4 / llama-3.1 families. Useful as a study set for LLM-as-judge robustness and a catalogue of steered-LLM failure modes. Note for LLM-as-judge research: the labels are themselves LLM-judge outputs — evaluating a judge (especially a Claude-family one) against this dataset partly measures consistency with Claude under this protocol, not independent human ground truth. AI is not a listed author; any errors are the author's. See README.md for the full taxonomy, schema, and labelling protocol.

提供机构:
Zenodo
创建时间:
2026-06-08
二维码
社区交流群
二维码
科研交流群
商业服务