FeatureBench
收藏资源简介:
# FeatureBench: Agent Coding Evaluation Benchmark ## Dataset Description FeatureBench is a comprehensive benchmark designed to evaluate AI agents' capabilities in end-to-end feature-level code generation. Unlike traditional benchmarks that focus on function-level or algorithm-specific tasks, FeatureBench challenges agents to implement complete features within real-world software projects. ### Key Characteristics - **Feature-Level Tasks**: Each task requires implementing a complete feature, including multiple functions, classes, and their interactions - **Real-World Codebases**: Tasks are derived from actual open-source projects, preserving the complexity and context of production code - **End-to-End Evaluation**: Agents must understand requirements, generate code, and pass comprehensive test suites - **Two Difficulty Levels**: - **Level 1 (lv1)**: Agents receive masked code with interface signatures and must implement the complete functionality - **Level 2 (lv2)**: Agents receive only test files and must implement both the interface and functionality from scratch ### Dataset Statistics - **Total Instances**: 330 - **full**: 200 instances - **lite**: 30 instances - **fast**: 100 instances - **Download Size**: 9.58 MB ## Dataset Structure Each instance in FeatureBench contains: - `instance_id`: Unique identifier for the task - `patch`: Git diff showing the implementation (Level 1) or empty string (Level 2) - `test_patch`: Git diff showing test file modifications - `FAIL_TO_PASS`: List of test files that must pass after implementation - `PASS_TO_PASS`: List of test files that must continue passing (Level 1 only) - `image_name`: Docker image containing the development environment - `repo`: Source repository (e.g., "owner/repo-name") - `base_commit`: Git commit hash of the base version - `problem_statement`: Detailed task description and requirements - `repo_settings`: Repository configuration settings as JSON string (from python.py) ## Usage ```python import json from datasets import load_dataset # Load a specific split dataset = load_dataset("LiberCoders/FeatureBench", split="lite") # Example: Access a task task = dataset[0] print(task['instance_id']) print(task['problem_statement']) # Parse repo_settings from JSON string repo_settings = json.loads(task['repo_settings']) print(repo_settings['repository']) print(repo_settings['base_image']) ```



