遇见数据集

AdithyaSK/repo2rlenv-pr-runtime

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

repo2rlenv-pr-runtime是一个用于强化学习(RL)和代码生成的Harbor格式数据集,由Repo2RLEnv工具生成,旨在将真实GitHub仓库转换为可验证的RL环境。数据集包含100个任务,基于13个开源GitHub仓库(如httpx、gin、mux、click、flask、werkzeug、requests、attrs、logrus、cobra、testify、typer、cli)的合并拉取请求或提交,通过pr_runtime管道提取并经过质量过滤,以移除指令中的信息泄露。每个任务提供自然语言指令、黄金补丁(oracle patch)、测试脚本和Dockerfile,奖励信号基于测试执行(包括FAIL_TO_PASS和PASS_TO_PASS测试的通过率),用于训练和评估大型语言模型(LLM)在代码修复和生成任务中的性能。数据集已通过oracle验证,所有任务均满足SWE-bench解析标准,并包含详细的验证指标(如resolved、command_resolved、eval_grade),适用于代码强化学习研究和基准测试。

repo2rlenv-pr-runtime is a Harbor-format dataset for reinforcement learning (RL) and code generation, generated by the Repo2RLEnv tool to turn real GitHub repositories into verifiable RL environments. It contains 100 tasks derived from merged pull requests or commits across 13 source GitHub repositories (e.g., httpx, gin, mux, click, flask, werkzeug, requests, attrs, logrus, cobra, testify, typer, cli), extracted via the pr_runtime pipeline with quality filters to strip information-leakage from instruction text. Each task includes a natural-language instruction, a gold patch (oracle), test scripts, and a Dockerfile, with a reward signal based on test execution (including F2P and P2P test pass rates) for training and evaluating large language models (LLMs) on code repair and generation tasks. The dataset is oracle-validated, with all tasks meeting SWE-bench resolution standards, and provides detailed validation metrics (e.g., resolved, command_resolved, eval_grade), suitable for code RL research and benchmarking.

提供机构:
AdithyaSK
二维码
社区交流群
二维码
科研交流群
商业服务