遇见数据集

synthiumjp/verbal-confidence-saturation

收藏
Hugging Face2026-04-27 更新2026-05-03 收录
官方服务:

资源简介:

Verbal Confidence Saturation数据集包含8,384个确定性试验,这些试验来自一个预注册研究,旨在测试3-9B指令调整的开源权重LLMs在最小引发下是否产生有效的口头置信度。八个开源模型在数字(0-100)和分类(10类)置信度引发下,使用贪婪解码方法回答了524个TriviaQA项目。所有七个指令模型在数字置信度上被心理测量有效性筛查分类为无效。平均上限率为91.7%。数据集中的每一行代表一个试验(模型×条件×项目),关键列包括模型ID、模型名称、条件、TriviaQA项目ID、问题文本、模型的原始响应、答案是否正确、提取的置信度值(0-1)、是否成功解析置信度、平均令牌对数概率、长度归一化的对数概率以及推理跟踪长度(仅M8模型)。

The Verbal Confidence Saturation Dataset consists of 8,384 deterministic trials from a pre-registered study testing whether 3–9B instruction-tuned open-weight LLMs produce valid verbal confidence under minimal elicitation. Eight open-weight models were administered 524 TriviaQA items under numeric (0–100) and categorical (10-class) confidence elicitation with greedy decoding. All seven instruct models were classified Invalid on numeric confidence by a psychometric validity screen. Mean ceiling rate: 91.7%. Each row in the dataset represents one trial (model × condition × item), with key columns including model ID, model name, condition, TriviaQA item ID, question text, models raw response, whether the answer was correct, extracted confidence value (0–1), whether confidence was successfully parsed, mean token logprobability, length-normalised logprobability, and reasoning trace length (M8 only).

提供机构:
synthiumjp
二维码
社区交流群
二维码
科研交流群
商业服务