stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t1.25_g8_run2_metrics
收藏资源简介:
这是一个编程任务数据集,用于分析代码质量和机器学习模型在代码生成或评估任务中的表现。数据集包含多个特征,如任务ID、入口点、可执行性、正确性、通过/失败的测试数量、测试运行时间、错误类型、Halstead复杂度度量(词汇量、长度、体积、难度、努力程度、时间)、圈复杂度、维护性指数、代码行数(LOC和SLOC)、注释百分比、类型标记比(TTR)、令牌字典、香农熵、平均预测熵、最大预测熵、定义函数数量以及入口点重复性。数据集基于训练分割,包含164个示例,总大小约243KB,下载大小约96KB。
This is a programming task dataset designed for analyzing code quality and the performance of machine learning models in code generation or evaluation tasks. The dataset includes features such as task ID, entry point, executability, correctness, number of tests passed/failed, test run time, error type, Halstead complexity measures (vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), token dictionary, Shannon entropy, mean predictive entropy, max predictive entropy, number of functions defined, and entry point repetition. It is based on a train split with 164 examples, total dataset size approximately 243KB, and download size approximately 96KB.




