SWE-Bench-Verified-Quick
收藏资源简介:
# SWE-Bench-Verified-Quick [](https://github.com/PrimeIntellect-ai/research-environments/tree/main/environments/swe/swebench_v1) Quick-eval subset of [`princeton-nlp/SWE-bench_Verified`](https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified) ([SWE-bench paper](https://arxiv.org/abs/2310.06770)): **468 / 500** Verified instances. Default dataset of the `swebench_v1` taskset; also supported by the v0 `mini_swe_agent_plus` environment. ## Changes vs upstream * **Latency subset only** — drops the slowest-running instances so a full benchmark pass finishes in under ~30 minutes at reasonable concurrency. No validation semantics; "Verified" in the name is OpenAI's human verification of the parent dataset. Schema and row content are unchanged. Upstream declares no dataset license; we mirror that and declare none. The data derives from public GitHub repositories under their own licenses; the SWE-bench project code is MIT. ## Splits | Split | Rows | |---|---:| | `test` | 468 | ## How to use Install the [`swebench_v1`](https://github.com/PrimeIntellect-ai/research-environments/tree/main/environments/swe/swebench_v1) taskset from [research-environments](https://github.com/PrimeIntellect-ai/research-environments), then run it end-to-end with [verifiers](https://github.com/PrimeIntellect-ai/verifiers): ```bash uv pip install --prerelease=allow "git+https://github.com/PrimeIntellect-ai/research-environments.git#subdirectory=environments/swe/swebench_v1" uv run eval --taskset.id swebench_v1 -m <your-model> -n 100 -r 4 ``` ## Original Dataset Card Snapshot of the [`princeton-nlp/SWE-bench_Verified`](https://huggingface.co/datasets/princeton-nlp/SWE-bench_Verified) card at card-build time — see the live card for updates. <details> <summary>Original <code>princeton-nlp/SWE-bench_Verified</code> dataset card</summary> **Dataset Summary** SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original SWE-bench dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? **Want to run inference now?** This dataset only contains the problem_statement (i.e. issue text) and the base_commit which represents the state of the codebase before the issue has been resolved. If you want to run inference using the "Oracle" or BM25 retrieval settings mentioned in the paper, consider the following datasets. princeton-nlp/SWE-bench_Lite_oracle princeton-nlp/SWE-bench_Lite_bm25_13K princeton-nlp/SWE-bench_Lite_bm25_27K **Supported Tasks and Leaderboards** SWE-bench proposes a new task: issue resolution provided a full repository and GitHub issue. The leaderboard can be found at www.swebench.com **Languages** The text of the dataset is primarily English, but we make no effort to filter or otherwise clean based on language type. **Dataset Structure** An example of a SWE-bench datum is as follows: ``` instance_id: (str) - A formatted instance identifier, usually as repo_owner__repo_name-PR-number. patch: (str) - The gold patch, the patch generated by the PR (minus test-related code), that resolved the issue. repo: (str) - The repository owner/name identifier from GitHub. base_commit: (str) - The commit hash of the repository representing the HEAD of the repository before the solution PR is applied. hints_text: (str) - Comments made on the issue prior to the creation of the solution PR’s first commit creation date. created_at: (str) - The creation date of the pull request. test_patch: (str) - A test-file patch that was contributed by the solution PR. problem_statement: (str) - The issue title and body. version: (str) - Installation version to use for running evaluation. environment_setup_commit: (str) - commit hash to use for environment setup and installation. FAIL_TO_PASS: (str) - A json list of strings that represent the set of tests resolved by the PR and tied to the issue resolution. PASS_TO_PASS: (str) - A json list of strings that represent tests that should pass before and after the PR application. ``` </details>



