遇见数据集

Supplementary material for FAULT: Failure Analysis and Understanding for Language-to-SPARQL Testing

收藏
Zenodo2026-05-07 更新2026-05-26 收录
官方服务:

资源简介:

Supplementary material for FAULT: Failure Analysis and Understanding for Language-to-SPARQL Testing Abstract: Converting natural language questions to SPARQL (NL2SPARQL) is a long-standing research challenge.To date, there are multiple NL2SPARQL benchmarks and open competitions.However, existing efforts have mainly focused on reporting aggregate accuracy metrics, which hinders further analysis of failure cases and error sources.In this work, we recognize the need for a structured framework to provide greater depth in existing and new NL2SPARQL benchmarks.We introduce the Failure Analysis and Understanding for Language-to-SPARQL Testing (FAULT) framework, an annotation framework for augmenting Natural-Language to SPARQL (NL2SPARQL) benchmarks enabling more fine-grained failure diagnosis.The framework allows for the identification of the system capabilities across two orthogonal dimensions: query complexity and question noise.This enables detailed analysis of system performance, identifying both specific query types that pose challenges and the linguistic noise most detrimental to accuracy.We apply this framework to augment an existing benchmark dataset (CK25 from the Text2SPARQL challenge) and create a new one based on the GPTKB knowledge graph.As a result, we produce semi-automatically annotated queries and a series of controlled test sets.Our observations reveal that recent LLM-based systems exhibit a tendency to hallucinate query elements, including predicates and entity IRIs, based on their names.Therefore, we also produce versions of the KGs with opaque IRIs to allow testing of a system's actual capability in understanding the schema of the targeted graph.We then present results from a representative system, demonstrating how our framework enables detailed analysis of the impact of lexicalized IRIs, linguistic noise in question formulation, and reliance on coherent query examples.This demonstrates the value of FAULT as a practical and effective resource for enabling more controlled and interpretable evaluation of NL2SPARQL systems. Github repository: https://github.com/UniVR-DH/nl2sparql-benchmark

FAULT:面向语言到SPARQL测试的故障分析与理解补充材料 摘要:将自然语言问题转换为SPARQL(NL2SPARQL)是一项长期存在的研究挑战。迄今为止,已涌现多款NL2SPARQL基准测试集与公开竞赛,但现有研究多聚焦于报告整体准确率指标,这极大阻碍了对故障案例与错误来源的深入分析。在本研究中,我们意识到亟需一套结构化框架,以深化对现有及新型NL2SPARQL基准测试集的分析。为此,我们提出了面向语言到SPARQL测试的故障分析与理解(Failure Analysis and Understanding for Language-to-SPARQL Testing,缩写FAULT)框架——这是一套用于扩充NL2SPARQL基准测试集的标注框架,可实现更细粒度的故障诊断。该框架可从两个正交维度识别系统能力:查询复杂度与问题噪声,借此可对系统性能展开精细化分析,既能定位具有挑战性的特定查询类型,也能找出对准确率损害最大的语言噪声类型。我们将该框架应用于扩充现有基准数据集(Text2SPARQL竞赛中的CK25),并基于GPTKB知识图谱构建了全新的基准测试集,最终生成了半自动化标注的查询语句与一系列可控测试集。我们的观测结果显示,近期基于大语言模型(LLM)的系统存在根据实体与谓词名称臆造查询元素(包括谓词与实体IRIs)的倾向。为此,我们还构建了带有不透明IRIs的知识图谱版本,以测试系统对目标图谱模式的真实理解能力。随后我们展示了一款代表性系统的测试结果,阐明了本框架可如何精细化分析词汇化IRIs、问题表述中的语言噪声以及对连贯查询示例的依赖所带来的影响。这充分证明了FAULT作为实用且高效的资源,可实现对NL2SPARQL系统更具可控性与可解释性的评估。 GitHub仓库:https://github.com/UniVR-DH/nl2sparql-benchmark

提供机构:
Zenodo
创建时间:
2026-05-07
二维码
社区交流群
二维码
科研交流群
商业服务