Fathom: Multi-Layer Coordinated Feature Suppression and Cognitive Autopsy for LLM Hallucination
收藏资源简介:
New in this version: Multi-layer coordinated feature suppression: Suppressing top 10 hallucination-associated SAE features across 22 layers (220 total) produces 33%% first-token changes on held-out TruthfulQA (KL=0.587, 73x larger than single-layer). First demonstration of coordinated SAE feature suppression crossing the argmax boundary. Feature decoding: Hallucination features at layer 11 encode structural/routing patterns (code tokens, BOS, templates), not semantic content. Hallucination is a computation pattern, not a factual error at the feature level. Intervention hierarchy: Single-layer suppression (0%% change), multi-layer (33%% change), best-of-N sampling (100%% CHS improvement). Direction of multi-layer changes is not consistently toward correctness (1 improved, 1 degraded, 8 unclear). Builds on v12 spectrum analysis (Gini AUC=0.685) and cognitive autopsy (layer 11 = 63%% fault epicenter). Related patents: US 64/020,489, US 64/021,113.



