ErrCalBench Benchmark Dataset for the paper "Self-Diagnostics and Rich Execution Feedback to Distill Models for Specific Software Engineering Tasks"
收藏官方服务:
资源简介:
The ErrCalBench dataset is a benchmark dataset released in conjunction with the paper “Self-Diagnostics and Rich Execution Feedback to Distill Modelsfor Specific Software Engineering Tasks” ErrCalBench is designed to evaluate and improve the self-diagnostic capabilities of task-optimized models during the execution of complex tasks. Unlike traditional single-output evaluations, this dataset emphasizes “rich execution feedback,” recording intermediate states, error types, and calibrated feedback during multi-step reasoning, tool invocation, and code execution
提供机构:
Zenodo创建时间:
2026-03-27



