遇见数据集

When do people rely on LLMs as they should? Measuring trust and trustworthiness on a single scale

收藏
Zenodo2026-04-29 更新2026-05-26 收录
官方服务:

资源简介:

Supporting data and analysis code for Teitelbaum et al. (2026), "When do people rely on LLMs as they should? Measuring trust and trustworthiness on a single scale." Abstract: People routinely consult Large Language Models (LLMs) for factual information, yet the accuracy of these systems varies dramatically across domains. How well do people calibrate their reliance to this uneven landscape of AI competence? We introduce a method for quantifying trust in an information source as its implicit evidential weight during incentivized belief updating, expressed on a common scale where 1 is defined as the evidential weight of a single human stranger's opinion. On the same scale, we estimate the ground-truth trustworthiness of LLM-generated information from observed error distributions. Across four preregistered studies, implicit trust was strongly domain-dependent. For public-opinion estimation, participants accorded ChatGPT roughly the same epistemic weight as a single human stranger (≈ 0.7), despite its true trustworthiness being somewhat higher (≈ 2.0). For numerical trivia, participants accorded ChatGPT far greater weight (≈ 23.5), though this still fell well below its actual trustworthiness when questions were posed in English to British participants (≈ 231). When the same questions were posed in Hebrew to Israeli participants, ChatGPT's trustworthiness dropped sharply (≈ 27.7), yet implicit trust remained comparable across the two groups. Explicit judgments, in contrast, substantially overestimated ChatGPT's epistemic weight in almost all cases and were only weakly related to implicit updating. Together, these findings suggest that people are sensitive to the broad contours of AI's 'jagged frontier' of competence, but do not calibrate to finer-grained variation within it.

提供机构:
Zenodo
创建时间:
2026-04-29
二维码
社区交流群
二维码
科研交流群
商业服务