ShahzebKhoso/local-code-arena-starcoder2_7b
收藏资源简介:
该数据集托管了原始评估指标、执行遥测日志和结构语法输出,这些数据来自对下一代StarCoder2 7B基础模型运行Mostly Basic Python Problems (MBPP)基准测试的结果。具体来说,它记录了现代中端原始基础权重在自动化对话评估工作流程中的行为动态,定义了原始代码能力和聊天后调优之间的明确界限。数据集包括500个任务(测试分割)的评估结果,涵盖功能通过率、平均生成速度等核心性能指标,并提供了详细的架构和特征模式,如任务ID、提示、标准参考、测试断言、模型元数据、原始生成、解析代码和评估指标等字段。
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the next-generation StarCoder2 7B base foundational model. This specific partition documents the behavioral dynamics of modern, mid-tier raw foundational weights inside automated conversational evaluation workflows, defining the clear boundaries between raw code capacity and chat post-tuning. The dataset includes evaluation results for 500 tasks (test split), covering core performance metrics such as functional pass rate and average generation speed, and provides detailed architecture and feature schemas, including fields like task_id, prompt, canonical_reference, test_assertions, model_metadata, raw_generation, parsed_code, and evaluation_metrics.




