stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.25_g10_run1_metrics
收藏资源简介:
该数据集是一个包含164个样本的训练集,用于代码分析或软件工程任务。每个样本包括任务ID、入口点、可执行性、正确性、测试通过/失败数量、测试运行时间(当前为空)、错误类型、Halstead软件度量(如词汇量、长度、体积、难度、努力和时间)、圈复杂度、可维护性指数、代码行数(总行数和源代码行数)、注释百分比、类型-标记比(TTR)、令牌字典、香农熵、预测熵(均值和最大值)、定义函数数量以及入口点重复性等特征。数据集大小为241,018字节,下载大小为96,592字节,适用于评估代码质量、复杂性或机器学习模型在编程相关应用中的性能。
This dataset is a training set containing 164 samples, designed for code analysis or software engineering tasks. Each sample includes features such as task ID, entry point, executability, correctness, number of tests passed/failed, test run time (currently null), error type, Halstead software metrics (e.g., vocabulary, length, volume, difficulty, effort, and time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), token dictionary, Shannon entropy, predictive entropy (mean and maximum), number of functions defined, and entry point repetition. The dataset size is 241,018 bytes with a download size of 96,592 bytes, suitable for evaluating code quality, complexity, or the performance of machine learning models in programming-related applications.




