Data and Analysis Materials for a Controlled Study of Fixed LLM Code-Generation Workflows
收藏资源简介:
This record contains a curated public release supporting a controlled study of fixed large-language-model code-generation workflows. It includes an analysis-ready table for 576 scheduled episodes across 48 selected HumanEval+ and MBPP+ tasks, four workflow bundles, and three outer repetitions; analysis scripts; workflow-stage prompt templates; environment and provenance summaries; compact result tables; and the internal planning and analysis-correction record. The release allows the reported analyses to be re-run from the sanitized episode table. It is not a complete end-to-end reproduction package: candidate programs, evaluator tests, raw model responses, rendered task-specific prompts, call-level ledgers, credentials, and service records are not redistributed. Consequently, users cannot independently rescore candidates or regenerate hosted-model outputs bit for bit.



