gemma-challenge/eval-prompts
收藏资源简介:
该数据集名为gemma-challenge/eval-prompts,是一个包含128个提示的集合,用于评估目的,特别是针对Gemma模型。这些提示从三个基准测试中采样:MMLU-Pro(来自TIGER-Lab/MMLU-Pro数据集的测试部分,57个提示)、GPQA Diamond(来自OpenAI simple-evals的gpqa_diamond.csv,57个提示)和AIME 2026(来自math-ai/aime26数据集,14个提示)。每个提示都是通过Inspect AI的inspect_evals工具在评估过程中捕获的,直接模拟发送给模型的消息格式,包括模板、答案选择格式和指令,没有手动编写。数据集提供了详细的列信息,如id、benchmark、source_dataset、messages、prompt、target和metadata,便于分析和使用。
A 128-prompt mix sampled from three benchmarks supported by Inspect AI / inspect_evals. Each prompt is rendered exactly as inspect_evals sends it to the model during evaluation — captured by running each task through Inspects mockllm/model and extracting the literal input messages (templates, answer-choice formatting, and instructions included). No prompt text was hand-written. The benchmarks include MMLU-Pro (57 prompts from TIGER-Lab/MMLU-Pro test set), GPQA Diamond (57 prompts from OpenAI simple-evals gpqa_diamond.csv), and AIME 2026 (14 prompts from math-ai/aime26 dataset). The dataset is designed for evaluation purposes, particularly for the Gemma model, and includes columns such as id, benchmark, source_dataset, messages, prompt, target, and metadata.




