stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.1_g10_run1_metrics
收藏资源简介:
该数据集是一个用于代码分析和评估的数据集,包含164个训练示例,每个示例具有多个特征,如任务ID、入口点、可执行性、正确性、测试通过和失败次数、测试运行时间(当前为空)、错误类型,以及一系列代码复杂度度量(如Halstead词汇量、长度、体积、难度、努力程度、时间,圈复杂度,可维护性指数),代码统计信息(如代码行数、源代码行数、注释百分比),文本特征(如类型标记比、香农熵、平均预测熵、最大预测熵),函数定义数量,以及入口点是否重复。数据集旨在支持代码质量、可维护性和性能的研究。
This dataset is designed for code analysis and evaluation, containing 164 training examples. Each example includes multiple features such as task ID, entry point, executability, correctness, number of tests passed and failed, test run time (currently null), error type, and a series of code complexity metrics (e.g., Halstead vocabulary, length, volume, difficulty, effort, time, cyclomatic complexity, maintainability index), code statistics (e.g., lines of code, source lines of code, comment percentage), text features (e.g., type-token ratio, Shannon entropy, mean predictive entropy, max predictive entropy), number of functions defined, and whether the entry point is repeated. The dataset aims to support research on code quality, maintainability, and performance.




