stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.1_g6_run2_metrics
收藏资源简介:
该数据集包含代码质量与复杂度相关的元数据,每条记录包含任务ID、入口点、是否可执行、是否正确、通过的测试数、失败的测试数、测试运行时间、错误类型、哈斯泰德度量(词汇量、长度、体积、难度、努力、时间)、圈复杂度、可维护性指数、代码行数、源代码行数、注释百分比、词类型比例(TTR)、令牌字典、香农熵、预测熵、定义的函数数、入口点是否重复等特征。数据集仅包含一个训练集,共164个样本,可能用于代码质量评估、复杂度分析或代码生成模型性能分析。
This dataset contains metadata related to code quality and complexity. Each record includes task ID, entry point, whether it is executable, whether it is correct, number of tests passed, number of tests failed, test run time, error type, Halstead metrics (vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (loc), source lines of code (sloc), comment percentage, token type ratio (TTR), token dictionary, Shannon entropy, predictive entropy, number of functions defined, and whether the entry point is repeated. The dataset consists of a single training split with 164 samples, likely used for code quality assessment, complexity analysis, or evaluating code generation model performance.




