stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t1.25_g3_run1_metrics
收藏资源简介:
该数据集包含代码分析相关的特征,用于评估编程任务的执行和复杂度。具体特征包括任务ID、入口点、可执行性、正确性、通过和失败的测试数量、测试运行时间(当前为空)、错误类型、Halstead复杂度指标(如词汇量、长度、体积、难度、努力和时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释比例、类型标记比(TTR)、令牌字典、香农熵、平均预测熵、最大预测熵、定义函数数量以及入口点是否重复。数据集包含一个训练分割,共有164个示例,适用于代码质量分析、机器学习模型训练或编程教育研究。
This dataset includes features related to code analysis for evaluating the execution and complexity of programming tasks. Specific features include task ID, entry point, executability, correctness, number of tests passed and failed, test run time (currently null), error type, Halstead complexity metrics (such as vocabulary, length, volume, difficulty, effort, and time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), token dictionary, Shannon entropy, mean predictive entropy, max predictive entropy, number of functions defined, and whether the entry point is repeated. The dataset contains a training split with 164 examples, suitable for code quality analysis, machine learning model training, or programming education research.




