neomatrix369/py-bug-trace-qwen3-6-27b-l1-rollouts
收藏资源简介:
该数据集是一个用于AI模型评估的结构化数据集,包含模型响应、奖励分数和性能指标。数据集包含15个训练示例,每个示例具有多个字段,如eval_id(评估ID)、model(模型名称)、level(级别)、example_id(示例ID)、rollout_number(展开编号)、trace_id(跟踪ID)、prompt(提示信息,包括角色、内容和工具调用)、completion(完成信息,包括角色、内容、推理内容和工具调用)、reward(奖励分数)、info(附加信息,如ID、难度、类别、时间统计、令牌使用情况、完成状态、截断状态、停止条件和指标)、exact_match_reward(精确匹配奖励)、latency_ms(延迟毫秒)和total_time(总时间)。info字段进一步包含timing(时间统计,包括开始时间、设置、生成、评分、模型、环境和总时间)和token_usage(令牌使用情况)等嵌套结构。数据集旨在支持多轮对话任务的分析和模型性能评估,重点关注响应质量、效率和准确性。
This dataset is a structured resource for AI model evaluation, encompassing model responses, reward scores, and performance metrics. It contains 15 training examples, each with multiple fields including: eval_id (evaluation ID), model (model name), level, example_id (example ID), rollout_number (rollout number), trace_id (trace ID), prompt (prompt information including role, content, and tool calls), completion (completion information including role, content, reasoning content, and tool calls), reward (reward score), info (additional information such as ID, difficulty, category, time statistics, token usage, completion status, truncation status, stop conditions, and metrics), exact_match_reward (exact match reward), latency_ms (latency in milliseconds), and total_time (total time). The info field further includes nested structures such as timing (time statistics covering start time, setup, generation, scoring, model, environment, and total time) and token_usage (token usage details). This dataset is designed to support analysis of multi-turn dialogue tasks and model performance evaluation, with a focus on response quality, efficiency, and accuracy.




