遇见数据集

anonymous-acl26/prompt-sensitivity-codegen

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

Anonymous Prompt Sensitivity Dataset是一个用于研究少样本代码生成中提示敏感性的数据集,包含约24万行数据。数据集覆盖多个模型(如claude-sonnet-4、gemini-2.5-flash、gpt-4o、llama-3.3-70b、qwen2.5-coder-3b)在两个基准测试(humaneval和mbpp)上的生成代码和评估结果。数据通过三个轴(order、phrasing、style)和k-shot值(0、1、2、3)进行组织,每行数据代表一个生成的样本,包括模型名称、基准测试、任务ID、问题类型、轴、变体、k-shot值、样本索引、生成代码、是否通过测试以及pass_at_1、pass_at_5、pass_at_10等评估指标。数据集旨在支持代码生成任务的提示敏感性分析,但未包含基准测试的原始提示、规范解决方案或测试代码,以确保匿名性和可复现性。

The Anonymous Prompt Sensitivity Dataset is a dataset for studying prompt sensitivity in few-shot code generation, containing approximately 240,000 rows. It covers multiple models (e.g., claude-sonnet-4, gemini-2.5-flash, gpt-4o, llama-3.3-70b, qwen2.5-coder-3b) on two benchmarks (humaneval and mbpp), with generated code and evaluation outcomes. The data is organized along three axes (order, phrasing, style) and k-shot values (0, 1, 2, 3). Each row represents one generated sample, including fields such as model, benchmark, task_id, problem_type, axis, variant, k_shot, sample_index, generated_code, passed, pass_at_1, pass_at_5, and pass_at_10. The dataset supports analysis of prompt sensitivity in code generation tasks but excludes benchmark source prompts, canonical solutions, or test code to maintain anonymity and reproducibility.

提供机构:
anonymous-acl26
二维码
社区交流群
二维码
科研交流群
商业服务