Sycophancy Scores on ELEPHANT: Five Response Dimensions Across Content, Style, and Framing
收藏资源简介:
Per-response sycophancy scores over the ELEPHANT advice scenarios (Cheng et al., ICLR 2026), scored on five dimensions — emotional acknowledgment, hedging, recommendation directness, counter-considerations, and endorsement — and measured across three input cues: content, style, and framing. The release contains 93,887 judgements, a crosswalk mapping each scenario back to its original ELEPHANT row or id, and the factual-preservation verdicts that define the 173-pair clean subset used in the framing analysis. Use it to analyze cue-level sycophancy on ELEPHANT without re-running any models, extend ELEPHANT's framing (AITA-NTA-FLIP) analysis with our preservation audit, or benchmark new models against these scores. It contains no raw text of any kind — no Reddit posts, generated scenarios, model responses, or judge rationales — only identifiers and integer scores. Every table in the accompanying paper (Decomposing LLM Advice Sycophancy, AACL-IJCNLP 2026) can be reproduced from these files. To recover the underlying situations, obtain ELEPHANT from its original source and join on the crosswalk. See README.md for the schema, joins, and a runnable reproduction snippet.



