RelationalFactQA
收藏资源简介:
RelationalFactQA数据集是一个用于评估大型语言模型(LLMs)从其内部参数知识中检索结构化、多记录表格输出的能力的新基准。该数据集包含696个问题、查询和答案三元组,涵盖九个知识领域,每个三元组都配有一个自然语言查询、等效的SQL语句和完全验证的黄金标准表格。该数据集旨在评估LLMs在生成结构化事实性知识方面的能力,特别是当输出维度(如属性或记录的数量)增加时。该数据集由手动和半自动生成,确保了多样性、可控的复杂性和覆盖范围。
The RelationalFactQA dataset is a novel benchmark for evaluating the capacity of large language models (LLMs) to retrieve structured, multi-record tabular outputs from their internal parametric knowledge. This dataset comprises 696 question-query-answer triples across nine knowledge domains, with each triple paired with a natural language query, an equivalent SQL statement, and a fully validated gold-standard table. This benchmark is designed to assess LLMs' capabilities in generating structured factual knowledge, especially as the dimensions of the output (such as the number of attributes or records) increase. It was generated via a combination of manual and semi-automated workflows to ensure diversity, controllable complexity, and comprehensive coverage.
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
基本信息
- 标题: RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
- 作者: Dario Satriani, Enzo Veltri, Donatello Santoro, Paolo Papotti
- 提交日期: 27 May 2025
- arXiv标识符: arXiv:2505.21409v1 [cs.CL]
- DOI: https://doi.org/10.48550/arXiv.2505.21409
摘要
- 研究背景: 大型语言模型(LLMs)的事实性是一个持续的挑战。当前基准测试通常评估简短的事实性答案,而忽略了从参数知识生成结构化、多记录表格输出的关键能力。
- 研究内容: 引入RelationalFactQA,一个专门设计用于评估结构化格式知识检索的新基准。该基准包含多样化的自然语言问题(与SQL配对)和黄金标准表格答案。
- 研究结果: 实验表明,即使是最先进的LLMs在生成关系输出方面也表现不佳,事实准确性不超过25%,且随着输出维度的增加性能显著下降。
主题分类
- 主要分类: Computation and Language (cs.CL)
- 次要分类:
- Artificial Intelligence (cs.AI)
- Databases (cs.DB)
相关资源
- PDF链接: View PDF
- TeX源码: TeX Source
- 其他格式: Other Formats
提交历史
- 版本1: [v1] Tue, 27 May 2025 16:33:38 UTC (304 KB)

- 1RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models意大利巴斯蒂亚塔大学 · 2025年



