stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t1.25_g9_run1_metrics
收藏资源简介:
该数据集包含编程任务的执行和代码度量数据,用于分析和评估代码质量。数据集包括任务ID、执行入口点、可执行性状态、正确性标志、测试通过和失败数量、测试运行时间(当前为空)、错误类型,以及多个代码复杂度指标(如Halstead词汇量、长度、体积、难度、努力程度和时间,圈复杂度,可维护性指数),同时涵盖代码行数(LOC和SLOC)、注释百分比、词汇多样性(TTR)、词汇字典、香农熵、预测熵均值和最大值,以及定义函数数量。数据分为训练集,包含164个样本,总大小约224KB,适用于代码质量评估、自动化测试或程序分析研究。
This dataset contains execution and code metric data for programming tasks, designed for analyzing and assessing code quality. It includes fields such as task ID, entry point, executability status, correctness flag, number of tests passed and failed, test run time (currently null), error type, and various code complexity metrics (e.g., Halstead vocabulary, length, volume, difficulty, effort, and time; cyclomatic complexity; maintainability index). Additionally, it covers lines of code (LOC and SLOC), comment percentage, vocabulary diversity (TTR), token dictionary, Shannon entropy, mean and maximum predictive entropy, and number of defined functions. The data is split into a training set with 164 examples and a total size of approximately 224KB, suitable for research in code quality evaluation, automated testing, or program analysis.




