stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.1_g6_run0_metrics
收藏资源简介:
该数据集包含164个训练样本,每个样本代表一个代码任务,具有多个特征字段,用于评估代码质量。特征包括任务ID、入口点、可执行性、正确性、测试通过和失败数、错误类型、Halstead复杂度指标(如词汇量、长度、体积、难度、努力程度、时间)、圈复杂度、维护性指数、代码行数(LOC和SLOC)、注释百分比、TTR(类型标记比)、令牌字典、香农熵、预测熵(均值和最大值)、定义函数数量以及入口点重复标志。数据集可能用于代码分析、机器学习模型训练或代码质量预测任务。
This dataset contains 164 training samples, each representing a code task with multiple feature fields for code quality assessment. Features include task ID, entry point, executability, correctness, tests passed and failed, error type, Halstead complexity metrics (e.g., vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, TTR (Type-Token Ratio), token dictionary, Shannon entropy, predictive entropy (mean and max), number of functions defined, and entry point repetition flag. The dataset is likely used for code analysis, machine learning model training, or code quality prediction tasks.




