遇见数据集

multilingual-swe-demo

收藏
魔搭社区2026-04-10 更新2026-07-15 收录
官方服务:

资源简介:

# Multilingual SWE-Bench Task Sample A sample package of a multilingual software engineering task benchmark dataset, designed to evaluate AI Agents' capabilities in code fixing and feature implementation on real open-source projects. Runs on the [Harbor](https://github.com/alexgshaw/harbor) evaluation framework. ## Overview This dataset contains **46 tasks** covering **9 programming languages** and **6 task types**, sourced from real open-source repository commits. Each task provides a Chinese `problem_statement` (problem description), requiring the Agent to modify code on top of a specified commit in the corresponding repository and pass predefined unit tests for verification. ## Directory Structure ``` demo/ ├── README.md # This file ├── sample_task.json # Metadata for all 46 tasks (instance_id, language, task type, test cases, etc.) ├── sample_images/ # Docker image tar archives (copy via cp_tar.py script, or prepare manually) │ └── <owner>__<repo>.tar # One pre-built Docker image per repository ├── sample_tasks/ # Individual directory for each task │ └── <owner>__<repo>__<hash>/ │ ├── task.toml # Task resource configuration (timeout, CPU, memory, storage limits) │ ├── instruction.md # Task description (detailed requirement document in Chinese RFC style) │ ├── environment/ │ │ ├── Dockerfile # Runtime environment definition (based on pre-built image) │ │ └── setup.sh # Environment initialization script (dependency installation, etc.) │ ├── tests/ │ │ ├── test.sh # Test execution and verification script (patch application, test run, reward calculation) │ │ └── config.json # Test configuration (log analysis script, f2p/p2p/p2f/f2f test classification) │ └── solution/ │ └── solve.sh # Reference solution (ground truth patch) └── test_script.sh # Quick verification script example ``` ## Quick Start ### Prerequisites - Python 3.10+ - Docker (for running isolated test environments) - Harbor evaluation framework ### Installation ```bash pip install harbor==0.1.36 ``` ### Loading Docker Images Each task runs in an isolated Docker container, so the corresponding repository image must be loaded first. Image tar archive names can be found in the `tar_path` field of each task entry in `sample_task.json`. ```bash # Load a single image docker load -i ./sample_images/<owner>__<repo>.tar # Example docker load -i ./sample_images/ajaxorg__ace.tar ``` ### Running Evaluation Use `test_script.sh` as a reference to run evaluations via the Harbor CLI: ```bash # Test Ground Truth (Oracle) — verify that the reference solution passes the tests python -m harbor.cli.main run \ --path ./sample_tasks/ajaxorg__ace__34e769c5 \ --agent oracle \ --model openrouter/anthropic/claude-opus-4.6 \ --jobs-dir ./jobs/ \ -k 1 # Test a custom Agent python -m harbor.cli.main run \ --path ./sample_tasks/ajaxorg__ace__34e769c5 \ --agent terminus-2 \ --model openrouter/anthropic/claude-opus-4.6 \ --jobs-dir ./jobs/ \ -k 1 ``` **Key Parameters:** | Parameter | Description | |-----------|-------------| | `--path` | Task directory path, pointing to a specific task under `sample_tasks/` | | `--agent` | Agent type: `oracle` uses the reference solution; other values use a custom Agent | | `--model` | LLM model identifier (via OpenRouter or other APIs) | | `--jobs-dir` | Evaluation results output directory | | `-k` | Number of attempts per task (k in pass@k) | Configure API keys before running: ```bash export OPENROUTER_API_KEY="your-api-key" export OPENROUTER_API_BASE="https://openrouter.ai/api/v1" ``` ## Data Format ### sample_task.json Each task entry contains the following fields: | Field | Type | Description | |-------|------|-------------| | `instance_id` | string | Unique task identifier, format: `<owner>__<repo>__<hash>` | | `task_type` | string | Task type code (see classification below) | | `task_type_reason` | string | Rationale for the task type classification | | `base_commit` | string | Base commit SHA of the target repository | | `language` | string | Primary programming language | | `repo_url` | string | Source repository GitHub URL | | `problem_statement` | string | Task description in Chinese | | `tar_path` | string | Corresponding Docker image tar filename | | `ut_results` | object | Unit test classification information | ### task.toml Defines resource limits for a single task: ```toml [verifier] timeout_sec = 3600 # Total timeout for verification phase [agent] timeout_sec = 3600 # Timeout for Agent working phase [environment] build_timeout_sec = 3600 # Docker build timeout cpus = 4 # Number of CPU cores memory = '16G' # Memory limit storage = '40G' # Storage limit ``` ## Task Type Distribution | Type | Full Name | Count | Description | |------|-----------|-------|-------------| | **BF** | Bug Fix | 20 | Fix code defects | | **FE** | Feature Enhancement | 12 | Feature enhancement / new feature implementation | | **FI** | Feature Implementation | 9 | Brand new feature implementation | | **RF** | Refactoring | 2 | Code refactoring | | **TG** | Test Generation | 2 | Test case generation | | **CD** | Code Documentation | 1 | Code documentation improvement | ## Language Distribution | Language | Count | |----------|-------| | JavaScript | 8 | | C | 7 | | Java | 7 | | Go | 5 | | PHP | 5 | | Ruby | 5 | | Swift | 4 | | Python | 3 | | C++ | 2 | ## Baseline Results | Model | pass@1 | |-------|--------| | Claude Opus 4.6 | 30.4% | ## Evaluation Pipeline The evaluation pipeline for each task is as follows: 1. **Environment Setup** — Launch a container from the pre-built Docker image and run `setup.sh` to install dependencies 2. **Agent Work** — The Agent reads the task description in `instruction.md`, modifies code inside the container, and generates a patch 3. **Patch Separation** — The Agent's changes are split into code changes (`code_changes.diff`) and test changes (`test_changes.diff`) 4. **Test Injection** — The Agent's test changes are discarded, and the predefined `test_patch.diff` is applied (ensuring test consistency) 5. **Test Execution** — Run the project's test suite and collect test logs 6. **Result Analysis** — Parse test logs via `analyze_test_logs.py`, compare against `f2p_tests` and `p2p_tests` to determine pass/fail 7. **Reward Calculation** — If all `f2p_tests` and `p2p_tests` pass, reward=1; otherwise reward=0

提供机构:
maas
创建时间:
2026-03-04
二维码
社区交流群
二维码
科研交流群
商业服务