kargaranamir/coercion
收藏资源简介:
该数据集名为Pressure-Coerced Self-Contradiction — mmlu,主要研究在高压环境下大型语言模型如何被迫产生自相矛盾的推理行为。数据集基于MMLU基准测试,使用meta-llama/Llama-3.3-70B-Instruct模型生成,包含2,052行数据。研究的关键指标包括基线正确率、任何强制成功率和信念崩溃率(BCR)等。数据集详细记录了模型在不同压力水平下的表现,包括原始问题、选择项、正确标签、压力水平、推理长度、归因方式等信息,以及模型在挑战中的对话记录和最终回答。
The dataset is named Pressure-Coerced Self-Contradiction — mmlu, which primarily investigates how large language models are coerced into producing self-contradictory reasoning under high-pressure conditions. Based on the MMLU benchmark, the dataset is generated using the meta-llama/Llama-3.3-70B-Instruct model and contains 2,052 rows. Key metrics include baseline correctness, any coercion success rate, and Belief Collapse Rate (BCR). The dataset meticulously documents model performance under various pressure levels, including original questions, choices, correct labels, pressure levels, reasoning length, attribution methods, as well as model conversation records and final answers in challenges.




