遇见数据集

noBuggie/AcslBench

收藏
Hugging Face2026-05-12 更新2026-05-31 收录
官方服务:

资源简介:

AcslBench 是一个面向 C 语言形式化规约合成的大规模已验证数据集,通过从 Rust/Verus 生态中迁移并严谨验证了超过 50 万条数据,解决了 C/ACSL 领域的“数据干旱”问题。数据集的核心亮点是引入了组合函数调用链,为评估大语言模型在复杂多函数推理场景下的形式化逻辑能力提供了可靠的“金标准”。数据集具有原生可验证性,所有样本均通过 Frama-C/WP 证明器验证,确保 ACSL 规约在数学上正确;包含高复杂度挑战,如最大循环嵌套深度达 18 层、验证条件超过 1,600 个的极端样本;并与工业级基准(如 Live-FMBench)在逻辑原语密度和代码模式上高度对齐,具有实用性。数据集分为两个子集:500k SFT 数据集,包含 50 万+样本,专为有监督微调和预训练设计,帮助模型学习 ACSL 语法和公理语义;AcslBench 组合评测基准,包含 495 个精选样本,按函数调用次数分为简单(≤2次调用)、中等(3次调用)和复杂(≥4次调用)三个等级,用于测试组合推理能力。

AcslBench is a large-scale verified dataset designed for formal C specification synthesis. It addresses the "data drought" in the C/ACSL domain by distilling over 500,000 verified instances from the Rust/Verus ecosystem. The core contribution is the introduction of Compositional Call Chains, providing a high-fidelity "gold standard" for evaluating the logical reasoning capabilities of Large Language Models in complex, multi-function scenarios. The dataset is verified by design, with every sample validated by the Frama-C/WP plugin to ensure mathematical soundness of ACSL specifications; it includes high complexity samples with up to 18-level nested loops and over 1,600 verification conditions, pushing the limits of modern formal reasoning; and it aligns closely with real-world industry benchmarks like Live-FMBench in terms of logical density and code patterns. The dataset consists of two subsets: the 500k SFT Dataset, with 500,000+ samples designed for Supervised Fine-Tuning and pre-training to learn ACSL syntax and axiomatic semantics; and the AcslBench Compositional Benchmark, a curated set of 495 samples categorized by Function Call Count into Simple Level (Call Count ≤ 2), Medium Level (Call Count == 3), and Complex Level (Call Count ≥ 4) to test compositional reasoning.

提供机构:
noBuggie
二维码
社区交流群
二维码
科研交流群
商业服务