stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.25_g5_run0_metrics
收藏资源简介:
该数据集是一个代码分析数据集,包含多个编程任务的评估指标和代码复杂度特征。每个样本代表一个任务,具有任务ID和入口点信息,并标注了可执行性、正确性、测试通过/失败数量、错误类型等评估结果。此外,数据集提供了丰富的代码度量指标,包括Halstead复杂度(词汇量、长度、体积、难度、努力程度、时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、TTR(类型标记比)、令牌字典、香农熵、预测熵(均值和最大值)以及定义函数数量。数据集用于分析代码质量、可维护性和测试性能,适用于机器学习或软件工程研究。数据集包含164个训练样本,文件大小约237KB。
This dataset is a code analysis dataset containing evaluation metrics and code complexity features for multiple programming tasks. Each sample represents a task with task ID and entry point information, annotated with executability, correctness, number of tests passed/failed, error type, and other evaluation results. Additionally, the dataset provides extensive code metrics, including Halstead complexity (vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, TTR (Type-Token Ratio), token dictionary, Shannon entropy, predictive entropy (mean and maximum), and number of defined functions. The dataset is used for analyzing code quality, maintainability, and testing performance, suitable for machine learning or software engineering research. It contains 164 training samples with a file size of approximately 237KB.




