遇见数据集

stefanocarrera/autophagycode_D_he_train-mercury_Qwen3-8B_strategy_trust_t0.75_g5_run0_metrics

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含代码任务执行结果和代码质量评估指标的数据,用于分析代码的可执行性、正确性、复杂度、可维护性等。具体特征包括任务ID、入口点、是否可执行、是否正确、测试通过和失败数量、错误类型,以及Halstead复杂度指标(如词汇量、长度、体积、难度、工作量、时间)、圈复杂度、可维护性指数、代码行数(LOC和SLOC)、注释百分比、类型标记比(TTR)、熵值(如香农熵、预测熵均值和最大值)、定义函数数量、入口点重复性等。数据集包含164个训练样本,适用于代码质量分析或机器学习任务。

This dataset contains data on code task execution results and code quality assessment metrics, used for analyzing code executability, correctness, complexity, maintainability, and more. Specific features include task ID, entry point, is_executable, is_correct, tests passed and failed counts, error type, as well as Halstead complexity metrics (e.g., vocabulary, length, volume, difficulty, effort, time), cyclomatic complexity, maintainability index, lines of code (LOC and SLOC), comment percentage, type-token ratio (TTR), entropy values (e.g., Shannon entropy, mean and max predictive entropy), number of defined functions, entry point repetition, etc. The dataset includes 164 training samples and is suitable for code quality analysis or machine learning tasks.

提供机构:
stefanocarrera
二维码
社区交流群
二维码
科研交流群
商业服务