Verifiable Linear Temporal Logic Benchmark (VLTL-Bench)
收藏资源简介:
VLTL-Bench是一个用于评估自然语言(NL)到线性时态逻辑(LTL)翻译系统的统一基准数据集,旨在衡量翻译的可验证性和正确性。该数据集包含三个独特的状态空间和数千个多样化的自然语言规范及其对应的时态逻辑规范,并提供样本跟踪以验证时态逻辑表达式。VLTL-Bench支持端到端评估,并提供每个步骤的真实值,以便研究人员改进和评估整个问题的不同子步骤。该数据集对于推动NL到LTL翻译领域的方法论研究具有重要意义。
VLTL-Bench is a unified benchmark dataset for evaluating natural language (NL) to linear temporal logic (LTL) translation systems, aiming to measure the verifiability and correctness of translations. This dataset includes three distinct state spaces, thousands of diverse natural language specifications and their corresponding temporal logic specifications, and provides sample traces to verify temporal logic expressions. VLTL-Bench supports end-to-end evaluation and provides ground truth for each step, enabling researchers to improve and assess different sub-steps of the overall problem. This dataset is of great significance for advancing methodological research in the field of NL-to-LTL translation.

- 1Verifiable Natural Language to Linear Temporal Logic Translation: A Benchmark Dataset and Evaluation SuiteUniversity of Florida, Gainesville, FL, USA; Florida International University, Miami, FL, USA · 2025年



