stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-4B_strategy_trust_t1.1_g5_run1_metrics
收藏资源简介:
该数据集是一个编程任务相关的数据集,包含164个训练样本,每个样本具有多项特征,如任务ID、入口点函数、可执行性、正确性、测试通过/失败数、错误类型、代码复杂度度量(Halstead词汇量、长度、体积、难度、努力度、时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、类型标记比(TTR)、标记字典、香农熵、预测熵均值/最大值、定义函数数量以及入口点重复性。这些特征用于分析和评估代码质量、可维护性和测试性能。数据集未提供具体应用场景或来源描述,但基于特征推断可能用于代码分析或机器学习任务。
This dataset is related to programming tasks, containing 164 training examples, each with multiple features such as task ID, entry point function, executability, correctness, number of tests passed/failed, error type, code complexity metrics (Halstead vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), token dictionary, Shannon entropy, mean/max predictive entropy, number of functions defined, and entry point repetition. These features are used for analyzing and evaluating code quality, maintainability, and test performance. The dataset does not provide specific application scenarios or source descriptions, but based on the features, it may be intended for code analysis or machine learning tasks.




