daksh76/prompt-sensitivity-codegen
收藏资源简介:
该数据集包含完整的生成代码输出和通过/失败结果,用于我们的提示敏感性研究,涵盖模型家族、基准测试、扰动轴和k-shot设置。具体来说,数据集有240000行数据,涉及模型包括claude-sonnet-4、gemini-2.5-flash、gpt-4o、llama-3.3-70b和qwen2.5-coder-3b;基准测试包括humaneval和mbpp;扰动轴包括order、phrasing和style;k-shot值包括0、1、2、3。数据集用于分析代码生成中提示的敏感性,帮助理解不同模型和设置下的性能变化。
This dataset contains the full generated-code outputs and pass/fail outcomes used in our prompt sensitivity study across model families, benchmarks, perturbation axes, and k-shot settings. Specifically, it includes 240,000 rows of data covering models such as claude-sonnet-4, gemini-2.5-flash, gpt-4o, llama-3.3-70b, and qwen2.5-coder-3b; benchmarks including humaneval and mbpp; axes like order, phrasing, and style; and k-shot values of 0, 1, 2, 3. It is designed to analyze prompt sensitivity in code generation, aiding in understanding performance variations across different models and settings.



