遇见数据集

OSReward

收藏
魔搭社区2026-09-05 更新2026-09-06 收录
官方服务:

资源简介:

# OSReward Benchmark OSReward evaluates whether a multimodal judge can determine if a computer-use agent completed a user's task. This release contains binary outcome labels only: `SUCCESS` and `FAIL`. The benchmark has two evaluation configurations: | Configuration | Trajectories | Unique task IDs | SUCCESS | FAIL | |---|---:|---:|---:|---:| | `full` | 1,019 | 656 | 440 | 579 | | `hard` | 284 | 257 | 86 | 198 | ## Load the dataset Load either benchmark configuration with `datasets`: ```python from datasets import load_dataset full = load_dataset("OS-Copilot/OSReward", "full", split="test") hard = load_dataset("OS-Copilot/OSReward", "hard", split="test") row = full[0] trajectory = row["trajectory"] screenshots = row["screenshots"] # Screenshot i is aligned with trajectory step i. image = screenshots[0] ``` Each record contains one complete trajectory. Original PNG/JPEG bytes are embedded in sharded Parquet files as a Hugging Face `List(Image)` column, so `load_dataset` handles downloading, caching, and image decoding without a TAR extraction step. The `full` and `hard` configurations are stored independently and retain the existing benchmark definitions. Reference evaluation code is maintained in the [OSReward GitHub repository](https://github.com/OS-Copilot/OSReward/tree/main/eval_pipeline). ## Canonical evaluation protocol The reference binary evaluation protocol uses: - the OSReward binary judge prompt; - full thought and action history; - last five screenshots; - red action-point markers when normalized coordinates are available; - temperature 0; - one verdict per trajectory: `Judge: SUCCESS` or `Judge: FAIL`. The primary metric for both configurations is **Balanced Accuracy**: ```text Balanced Accuracy = (SUCCESS Recall + FAIL Recall) / 2 ``` Also report Accuracy, SUCCESS Recall, FAIL Recall, and Coverage. Missing, API-error, or unparseable outputs count as incorrect. This strict error policy prevents a judge from improving its score by abstaining. ## Record format Each Parquet row contains one trajectory: ```json { "trace_id": "...", "task_id": "...", "platform": "Web", "agent": "...", "instruction": "...", "frame_semantics": "pre", "trajectory_length": 10, "trajectory": [ { "step_index": 0, "screenshot_path": "../../screenshots/<trace_id>/step_0000.png", "thought": "...", "action": "...", "coordinate": [500, 500] } ], "screenshots": ["<Image>", "..."], "human_label": "SUCCESS" } ``` `screenshots` and `trajectory` always have the same length. The value at `screenshots[i]` is the frame for `trajectory[i]`. The legacy `screenshot_path` string is retained as provenance, but consumers should read the embedded `screenshots` column. Coordinates are normalized to `[0, 1000]`. A coordinate is `null` when an action has no single point to mark. `frame_semantics` is `pre`: screenshot *i* shows the state before action *i*. The effect of the last action may therefore be absent from the final frame. Eighty-one steps in Full have `screenshot_path: null` and an aligned `screenshots` value of `None` because the source release provides no image reference for those positions. Eighty of these steps belong to one FAIL trajectory, which is evaluated from its complete text history without visual input; one other trajectory is missing a single frame. The evaluator skips unavailable selected frames but retains every step's text history. Hard contains no missing images. ## Data integrity - 1,019 unique trace IDs in Full; - 284 unique trace IDs in Hard; - Hard is byte-for-byte identical to the corresponding Full records; - 30,588 referenced screenshots; - contiguous zero-based step indices; - no binary label or screenshot-path mismatch between release views; - every embedded image is byte-identical to its source screenshot; - `screenshots[i]` is aligned with `trajectory[i]` in every record. The release was checked for schema consistency, subset identity, screenshot alignment, source-byte identity, and image decodability before publication. The sharded Parquet files are the canonical and only published data artifacts. They contain both trajectory metadata and the aligned screenshot bytes. ## Citation ```bibtex @article{sun2026osreward, title={OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models}, author={Qiushi Sun and Kanzhi Cheng and Yian Wang and Bowen Yang and Hang Yan and Liheng Chen and Fangzhi Xu and Zichen Ding and Nuo Chen and Jialin Cao and Xingdong Gong and Zehao Li and Kaiming Jin and Xinfeng Yuan and Zhoumianze Liu and Jingyang Gong and Zhangyue Yin and Jiahui Gao and Zhiyong Wu and Tianbao Xie and Jianbing Zhang and Ben Kao and Lingpeng Kong}, journal={arXiv preprint arXiv:2607.28609}, year={2026} } ```

提供机构:
maas
创建时间:
2026-07-31
二维码
社区交流群
二维码
科研交流群
商业服务