stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t0.75_g8_run0_metrics
收藏资源简介:
该数据集是一个用于代码分析和评估的数据集,包含多个特征字段,如任务ID、入口点函数、代码可执行性、正确性、测试通过和失败数量、错误类型,以及一系列代码复杂度度量指标(包括Halstead词汇量、长度、体积、难度、努力度和时间,圈复杂度,可维护性指数,代码行数,注释百分比等)。此外,还包含词汇多样性、熵值和预测熵等统计信息。数据集适用于机器学习模型训练、代码质量评估或软件工程研究,支持对代码片段进行多维度分析和分类任务。
This dataset is designed for code analysis and evaluation, featuring multiple attributes such as task ID, entry point function, code executability, correctness, number of tests passed and failed, error types, and a range of code complexity metrics (including Halstead vocabulary, length, volume, difficulty, effort, and time, cyclomatic complexity, maintainability index, lines of code, comment percentage, etc.). It also includes statistical measures like token type ratio, Shannon entropy, and predictive entropy. The dataset is suitable for machine learning model training, code quality assessment, or software engineering research, enabling multi-dimensional analysis and classification of code snippets.




