stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.25_g1_run2_metrics
收藏资源简介:
该数据集包含代码分析相关的数据,用于评估编程任务的执行和代码质量。它提供了164个训练样本,每个样本包含任务ID、入口点函数、可执行性标志、正确性标志、测试通过和失败数量、错误类型,以及多种代码复杂度指标(如Halstead词汇量、长度、体积、难度、努力程度和时间,圈复杂度,可维护性指数),还包括代码行数(LOC和SLOC)、注释比例、词汇多样性(TTR)、香农熵和预测熵等特征。数据集可能用于机器学习模型训练,以分析代码性能、错误检测或代码质量评估。
This dataset contains code analysis-related data for evaluating the execution and code quality of programming tasks. It provides 164 training samples, each including task ID, entry point function, executability flag, correctness flag, number of tests passed and failed, error type, and various code complexity metrics (such as Halstead vocabulary, length, volume, difficulty, effort, and time, cyclomatic complexity, maintainability index), as well as lines of code (LOC and SLOC), comment percentage, lexical diversity (TTR), Shannon entropy, and predictive entropy. The dataset is likely used for machine learning model training to analyze code performance, error detection, or code quality assessment.




