stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.25_g4_run2_metrics
收藏资源简介:
该数据集包含164个训练样本,用于代码质量评估和分析。每个样本代表一个代码任务,特征包括任务ID、入口点、可执行性、正确性、通过和失败的测试数量、测试运行时间(当前为空)、错误类型、Halstead复杂度指标(词汇量、长度、体积、难度、努力程度、时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、类型标记比(TTR)、令牌字典、香农熵、平均和最大预测熵、定义函数数量以及入口点是否重复。数据集旨在支持代码执行性能、错误检测和代码度量研究。
This dataset contains 164 training samples for code quality evaluation and analysis. Each sample represents a code task, with features including task ID, entry point, executability, correctness, number of passed and failed tests, test runtime (currently empty), error type, Halstead complexity metrics (vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), token dictionary, Shannon entropy, average and maximum prediction entropy, number of defined functions, and whether the entry point is duplicated. This dataset is designed to support research on code execution performance, error detection, and code metrics.




