FailDiag: Selected Program Failure Diagnosis Benchmark
收藏资源简介:
FailDiag is a benchmark for evaluating whether an LLM diagnosis identifies the defect that explains a selected program failure rather than merely identifying any defect in the program. The release contains 600 real C++ student-program cases from 110 programming tasks, balanced across Programming Fundamentals and Data Structures and Algorithms. Each case includes the faulty program, selected expected-versus-observed failure evidence, same-lineage repair evidence, and structured diagnostic references. The public release also provides the exact original task text used in the reported experiments, an English-only task rendering for accessibility, a condition-safe loader for Code-Only, Failure-Informed, and Repair-Assisted evaluation, schemas, validation reports, and audit documentation. The archived files correspond to Hugging Face dataset revision:0aa5c7a19daad4595fbc3dc43ffa03cf79a52861 Public dataset:https://huggingface.co/datasets/huytran1499/FailDiag Associated manuscript:FailDiag: Evaluating LLMs for Diagnosis of Selected Program Failures



