neomatrix369/py-bug-trace-qwen3-5-35b-a3b-l1-rollouts
收藏资源简介:
该数据集是一个用于评估模型性能的结构化数据集,包含多个特征字段,主要用于记录和分析模型在任务中的表现。数据集包括评估ID(eval_id)、模型名称(model)、难度级别(level)、示例ID(example_id)、滚动编号(rollout_number)、追踪ID(trace_id)、提示信息(prompt,包含角色、内容等)、完成响应(completion,包含角色、内容、推理内容等)、奖励分数(reward)、详细信息(info,如ID、难度、类别、时间记录、令牌使用情况、完成状态、截断状态、停止条件、指标如精确匹配奖励和回合数)、精确匹配奖励(exact_match_reward)、延迟毫秒(latency_ms)和总时间(total_time)。时间记录部分详细记录了开始时间、设置、生成、评分、模型、环境等各个阶段的时间跨度。数据集仅包含训练分割(train),有15个示例,总大小约为44KB,下载大小约为74KB。推断该数据集可能用于机器学习或人工智能模型的基准测试、性能评估或交互式任务分析,侧重于模型响应质量、效率和时序指标。
This dataset is a structured dataset for evaluating model performance, containing multiple feature fields primarily used to record and analyze model performance in tasks. The dataset includes evaluation ID (eval_id), model name (model), difficulty level (level), example ID (example_id), rollout number (rollout_number), trace ID (trace_id), prompt information (prompt, including role, content, etc.), completion response (completion, including role, content, reasoning content, etc.), reward score (reward), detailed information (info, such as ID, difficulty, category, timing records, token usage, completion status, truncation status, stop condition, metrics like exact match reward and number of turns), exact match reward (exact_match_reward), latency in milliseconds (latency_ms), and total time (total_time). The timing section records detailed time spans for various stages including start time, setup, generation, scoring, model, environment, etc. The dataset only includes a training split (train) with 15 examples, total size approximately 44KB, and download size approximately 74KB. It is inferred that this dataset may be used for benchmarking, performance evaluation, or interactive task analysis of machine learning or AI models, focusing on model response quality, efficiency, and timing metrics.




