stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.1_g4_run2_metrics
收藏资源简介:
该数据集包含164个训练样本,用于代码执行和评估任务。每个样本具有多个特征,包括任务ID、入口点、可执行性、正确性、测试通过和失败数量、错误类型,以及代码复杂度指标(如halstead词汇量、长度、体积、难度、努力和时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、TTR(类型-标记比)、令牌字典、香农熵、平均和最大预测熵、定义函数数量、入口点重复性等。数据集旨在支持代码质量分析、自动化测试和机器学习模型训练。
This dataset contains 164 training samples for code execution and evaluation tasks. Each sample includes multiple features such as task ID, entry point, executability, correctness, number of tests passed and failed, error type, and code complexity metrics (e.g., Halstead vocabulary, length, volume, difficulty, effort, and time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, TTR (Type-Token Ratio), token dictionary, Shannon entropy, mean and maximum predictive entropy, number of functions defined, and entry point repetition. The dataset is designed to support code quality analysis, automated testing, and machine learning model training.




