遇见数据集

Supplementary material for FAULT: Failure Analysis and Understanding for Language-to-SPARQL Testing

收藏
Zenodo2026-05-18 更新2026-05-26 收录
官方服务:

资源简介:

Supplementary material for FAULT: Failure Analysis and Understanding for Language-to-SPARQL Testing Abstract: Converting natural language questions to SPARQL (NL2SPARQL) is a long-standing research challenge.To date, there are multiple NL2SPARQL benchmarks and open competitions.However, existing efforts have mainly focused on reporting aggregate accuracy metrics, which hinders further analysis of failure cases and error sources.In this work, we recognize the need for a structured framework to provide greater depth in existing and new NL2SPARQL benchmarks.We introduce the Failure Analysis and Understanding for Language-to-SPARQL Testing (FAULT) framework, an annotation framework for augmenting Natural-Language to SPARQL (NL2SPARQL) benchmarks enabling more fine-grained failure diagnosis.The framework allows for the identification of the system capabilities across two orthogonal dimensions: query complexity and question noise.This enables detailed analysis of system performance, identifying both specific query types that pose challenges and the linguistic noise most detrimental to accuracy.We apply this framework to augment an existing benchmark dataset (CK25 from the Text2SPARQL challenge) and create a new one based on the GPTKB knowledge graph.As a result, we produce semi-automatically annotated queries and a series of controlled test sets.Our observations reveal that recent LLM-based systems exhibit a tendency to hallucinate query elements, including predicates and entity IRIs, based on their names.Therefore, we also produce versions of the KGs with opaque IRIs to allow testing of a system's actual capability in understanding the schema of the targeted graph.We then present results from a representative system, demonstrating how our framework enables detailed analysis of the impact of lexicalized IRIs, linguistic noise in question formulation, and reliance on coherent query examples.This demonstrates the value of FAULT as a practical and effective resource for enabling more controlled and interpretable evaluation of NL2SPARQL systems. Github repository: https://github.com/UniVR-DH/nl2sparql-benchmark

提供机构:
Zenodo
创建时间:
2026-05-18
二维码
社区交流群
二维码
科研交流群
商业服务