stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t1.25_g2_run2_metrics
收藏资源简介:
该数据集包含164个编程任务样本,每个样本记录任务ID、入口点、可执行性、正确性、测试通过/失败数、错误类型等基本属性,并提供了丰富的代码质量评估指标,包括Halstead复杂度度量(词汇量、长度、体积、难度、工作量、时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、词汇多样性(TTR)、香农熵、预测熵均值与最大值、定义函数数量及入口点重复标志。数据集适用于代码质量分析、自动化测试评估、编程教育辅助和软件工程研究。
This dataset contains 164 programming task samples, each recording basic attributes such as task ID, entry point, executability, correctness, tests passed/failed, and error type. It provides comprehensive code quality assessment metrics including Halstead complexity measures (vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, lexical diversity (TTR), Shannon entropy, mean and maximum predictive entropy, number of defined functions, and entry point repetition flag. The dataset is suitable for code quality analysis, automated testing evaluation, programming education support, and software engineering research.




