stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t1.25_g10_run0_metrics
收藏资源简介:
该数据集包含164个训练样本,用于代码分析或编程任务评估。每个样本代表一个代码任务,特征包括任务ID、入口点函数、可执行性、正确性、测试通过和失败数量、错误类型,以及多种代码度量指标,如Halstead复杂度(词汇量、长度、体积、难度、努力程度、时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、词汇多样性(TTR)、香农熵、预测熵(均值和最大值)、定义函数数量,以及入口点是否重复。数据集旨在支持代码质量评估、性能分析或机器学习模型训练,适用于编程语言处理或软件工程研究。
This dataset contains 164 training samples for code analysis or programming task evaluation. Each sample represents a code task, with features including task ID, entry point function, executability, correctness, number of tests passed and failed, error type, and various code metrics such as Halstead complexity (vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), Shannon entropy, predictive entropy (mean and maximum), number of functions defined, and whether the entry point is repeated. The dataset is designed to support code quality assessment, performance analysis, or machine learning model training, suitable for programming language processing or software engineering research.



