遇见数据集

Sycophancy Scores on ELEPHANT: Five Response Dimensions Across Content, Style, and Framing

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

Per-response sycophancy scores over the ELEPHANT advice scenarios (Cheng et al., ICLR 2026), scored on five dimensions — emotional acknowledgment, hedging, recommendation directness, counter-considerations, and endorsement — and measured across three input cues: content, style, and framing. The release contains 93,887 judgements, a crosswalk mapping each scenario back to its original ELEPHANT row or id, and the factual-preservation verdicts that define the 173-pair clean subset used in the framing analysis. Use it to analyze cue-level sycophancy on ELEPHANT without re-running any models, extend ELEPHANT's framing (AITA-NTA-FLIP) analysis with our preservation audit, or benchmark new models against these scores. It contains no raw text of any kind — no Reddit posts, generated scenarios, model responses, or judge rationales — only identifiers and integer scores. Every table in the accompanying paper (Decomposing LLM Advice Sycophancy, AACL-IJCNLP 2026) can be reproduced from these files. To recover the underlying situations, obtain ELEPHANT from its original source and join on the crosswalk. See README.md for the schema, joins, and a runnable reproduction snippet.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务