stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t1.25_g9_run2_metrics
收藏资源简介:
该数据集包含一系列代码任务的相关数据,每个条目由任务ID、入口点、可执行性标志、正确性标志、通过和失败的测试数量等字段组成,并提供了详细的代码度量指标,如Halstead度量(包括词汇量、长度、体积、难度、努力和时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比等。此外,还包含文本分析指标,如类型-标记比(TTR)、香农熵和预测熵,用于评估代码的复杂性和信息内容。数据集适用于软件工程研究、代码质量评估、NLP应用或机器学习任务,旨在分析代码特性与执行结果之间的关系。
This dataset includes data related to a series of code tasks, with each entry consisting of fields such as task ID, entry point, executability flag, correctness flag, number of tests passed and failed, and provides detailed code metrics like Halstead measures (including vocabulary, length, volume, difficulty, effort, and time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, etc. It also contains text analysis metrics such as type-token ratio (TTR), Shannon entropy, and predictive entropy for assessing code complexity and information content. The dataset is suitable for software engineering research, code quality evaluation, NLP applications, or machine learning tasks, aiming to analyze the relationship between code characteristics and execution outcomes.




