遇见数据集

TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability

收藏
The Canadian Dataverse Repository2024-01-01 更新2026-04-17 收录
官方服务:

资源简介:

Large Language Model (LLM) evaluation is currently one of the most important areas of research, with existing benchmarks proving to be insufficient and not completely representative of LLMs' various capabilities. We present a curated collection of challenging statements on sensitive topics for LLM benchmarking called TruthEval. These statements were curated by hand and contain known truth values. The categories were chosen to distinguish LLMs' abilities from their stochastic nature. Details of collection method and use cases can be found in this paper: <a href="https://arxiv.org/abs/2406.01855">TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability</a>

提供机构:
University of Waterloo
创建时间:
2024-01-01
二维码
社区交流群
二维码
科研交流群
商业服务