遇见数据集

ucberkeley-dlab/fragility-moral-judgment-llms

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是FAccT论文《大型语言模型中道德判断的脆弱性》的配套数据集,用于研究大型语言模型(LLM)在道德判断上的稳定性。数据集包含从Reddit的r/AmItheAsshole子版块收集的道德困境、社区标签分布以及多个LLM模型的判决结果(包括解释和推理轨迹)。通过应用最小化且道德无关的扰动(如删除句子、改变琐碎细节等),研究LLM对同一道德困境的判断是否稳定,并探讨不同协议和推理链对稳定性的影响。数据集分为多个配置:dilemmas(源困境数据)、model_verdicts(主要评估数据,包含模型判决)、protocol_variations(不同协议下的评估)、reasoning_traces(推理能力模型的推理链)、verification_annotations(对推理轨迹的LLM作为法官的注释)和entropy_baseline(基准熵值数据)。数据用于评估LLM道德推理的鲁棒性和公平性,但不应作为道德判断的 ground truth 或用于训练道德仲裁模型。数据集在CC BY 4.0许可证下发布,原始Reddit文本受其作者版权保护。

Companion dataset for the FAccT paper Fragility of Moral Judgment in Large Language Models. Contains moral dilemmas from the r/AmItheAsshole subreddit, community label distributions, and per-model verdicts (including explanations and reasoning traces) used to study the stability of LLM moral judgments under minimal, morally-irrelevant perturbations. The dataset investigates how LLM judgments vary across different perturbation types (e.g., robustness and framing perturbations) and protocols (e.g., verdict-first vs. explanation-first). It includes multiple configs: dilemmas (source dilemmas), model_verdicts (main evaluation data), protocol_variations (protocol-based evaluations), reasoning_traces (reasoning chains from capable models), verification_annotations (LLM-as-judge annotations), and entropy_baseline (baseline entropy measures). Intended for research on LLM robustness and fairness in moral reasoning, but not as ground truth for moral judgments. Released under CC BY 4.0, with original Reddit text subject to authors copyright.

提供机构:
ucberkeley-dlab
二维码
社区交流群
二维码
科研交流群
商业服务