遇见数据集

proxectonos/veritasqa_gl

收藏
Hugging Face2026-04-27 更新2026-05-10 收录
官方服务:

资源简介:

VeritasQA GL是VeritasQA的加利西亚语版本,这是一个用于评估问答系统和语言模型真实性的基准。该基准旨在测试模型是否复制常见误解和虚假信息,而非提供真实答案。数据集包含360行数据,每行包括问题、最佳答案、可接受正确和错误答案集、类别和来源字段。它适用于真实性评估、开放域问答、多语言LLM评估以及基准测试对常见误解和虚假信息的抵抗能力。语言为加利西亚语,但VeritasQA基准也支持西班牙语、加泰罗尼亚语和英语。数据集创建围绕独立于特定上下文、国家或近期事件的问题,针对广泛误解和虚假信息,旨在提供比早期基准更可转移和持久的评估资源。

VeritasQA GL is the Galician version of VeritasQA, a benchmark for evaluating the truthfulness of question-answering systems and language models. The benchmark is designed to test whether models reproduce common misconceptions and falsehoods rather than giving truthful answers. The dataset contains 360 rows, each including a question, a best answer, sets of acceptable correct and incorrect answers, a category, and a source field. It is suitable for truthfulness evaluation, open-domain question answering, evaluation of multilingual LLMs, and benchmarking resistance to common misconceptions and falsehoods. The language is Galician, but the VeritasQA benchmark is also available in Spanish, Catalan, and English. The dataset is built around questions largely independent of specific contexts, countries, or recent events, targeting widespread misconceptions and falsehoods to provide a more transferable and durable evaluation resource than earlier benchmarks.

提供机构:
proxectonos
二维码
社区交流群
二维码
科研交流群
商业服务