stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t0.75_g10_run1_metrics
收藏资源简介:
该数据集包含代码分析或编程任务相关的数据,具体特征包括任务ID、入口点函数、可执行性状态、正确性标记、测试通过和失败数量、测试运行时间、错误类型,以及多种代码度量指标如Halstead复杂度(词汇量、长度、体积、难度、努力程度、时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、类型标记比(TTR)、令牌字典、香农熵、平均预测熵、最大预测熵、定义函数数量和入口点重复标记。数据集分为训练集,包含164个示例,总大小约230KB。
This dataset contains data related to code analysis or programming tasks, with features including task ID, entry point function, executability status, correctness flag, number of tests passed and failed, test run time, error type, and various code metrics such as Halstead complexity (vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), token dictionary, Shannon entropy, mean predictive entropy, max predictive entropy, number of functions defined, and entry point repetition flag. The dataset is split into a training set with 164 examples and a total size of approximately 230KB.




