stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.25_g6_run1_metrics
收藏资源简介:
该数据集包含编程任务相关的数据,用于分析代码执行和复杂度。特征包括任务ID、入口点、可执行性、正确性、通过/失败的测试数量、测试运行时间(当前为空)、错误类型、Halstead复杂度指标(如词汇量、长度、体积、难度、努力度和时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、TTR(类型标记比)、令牌字典、香农熵、预测熵(均值和最大值)、定义函数数量以及入口点重复标志。数据集适用于代码质量评估、软件工程研究或机器学习任务,如代码生成或缺陷检测。
This dataset contains data related to programming tasks, designed for analyzing code execution and complexity. Features include task ID, entry point, executability, correctness, number of tests passed/failed, test run time (currently null), error type, Halstead complexity metrics (e.g., vocabulary, length, volume, difficulty, effort, and time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, TTR (Type-Token Ratio), token dictionary, Shannon entropy, predictive entropy (mean and max), number of functions defined, and entry point repetition flag. The dataset is suitable for code quality assessment, software engineering research, or machine learning tasks such as code generation or defect detection.




