stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.1_g2_run2_metrics
收藏资源简介:
该数据集是一个用于代码质量评估和程序分析的数据集,包含164个训练示例。数据集提供了丰富的代码特征,包括任务ID、入口点、可执行性、正确性、测试通过和失败数量、测试运行时间、错误类型、Halstead复杂度指标(如词汇量、长度、体积、难度、努力程度和时间)、圈复杂度、可维护性指数、代码行数(总行数和源代码行数)、注释百分比、类型标记比(TTR)、令牌字典、香农熵、预测熵(均值和最大值)、定义函数数量以及入口点重复性。这些特征旨在支持代码分析、测试验证和编程任务评估,适用于机器学习和自然语言处理在代码生成和质量检测中的应用。
This dataset is designed for code quality assessment and program analysis, containing 164 training examples. It provides extensive code features, including task ID, entry point, executability, correctness, number of tests passed and failed, test run time, error type, Halstead complexity metrics (such as vocabulary, length, volume, difficulty, effort, and time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), token dictionary, Shannon entropy, predictive entropy (mean and maximum), number of functions defined, and entry point repetition. These features are intended to support code analysis, test verification, and programming task evaluation, suitable for applications in machine learning and natural language processing for code generation and quality detection.




