遇见数据集

leibni/truthful_qa

收藏
Hugging Face2026-04-26 更新2026-05-03 收录
官方服务:

资源简介:

TruthfulQA是一个用于衡量语言模型在生成问题答案时是否真实的基准测试。该基准包含817个问题,涵盖38个类别,包括健康、法律、金融和政治。问题设计巧妙,使得一些人会因错误信念或误解而给出错误答案。为了表现良好,模型必须避免生成从模仿人类文本中学到的错误答案。

TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.

提供机构:
leibni
二维码
社区交流群
二维码
科研交流群
商业服务