遇见数据集

Robbie's Razor Benchmark Results v0.1.0: Preregistered Bounded Evaluation Dataset

收藏
Zenodo2026-08-17 更新2026-08-20 收录
官方服务:

资源简介:

Robbie’s Razor Benchmark Results v0.1.0 is a preregistered, first-party bounded evaluation dataset produced under protocol RR-BRP-0.1.0 using Robbie’s Razor Benchmarks v0.2.0. The evaluation covers three OpenAI models—gpt-5.6-luna, gpt-5.6-terra, and gpt-5.6-sol—under two matched conditions, API-C0 and API-R1. It includes four evaluator-defined cases, three repetitions per case and condition, 72 completed live API requests, 18 evaluation reports, six aggregate model-condition groups, and zero recorded API errors. API response storage was disabled using store: false. Within this documented task set, observed accuracy was 0.75 under API-C0 and 1.00 under API-R1 for gpt-5.6-luna; 0.50 under API-C0 and 1.00 under API-R1 for gpt-5.6-terra; and 1.00 under both conditions for gpt-5.6-sol. These results apply only to the models, prompts, cases, conditions, repetitions, evaluator, API configuration, and execution environment preserved in this dataset. The package includes raw execution records, per-run outputs, evaluator reports, aggregate summaries, frozen benchmark inputs, environment and model metadata, authorization records, derived metrics, deviation records, checksums, and a machine-readable results manifest. The archived ZIP has SHA-256 checksum 93010c38e26292c30b337e23128727bb0cdf9640cbd2b355616310b6fcf84e4d. The synthetic memory-gate stage initially encountered a Python module-resolution error after all 72 live requests had completed. It was recovered separately without repeating any live model request. The event is documented as RR-BRP-0.1.0-DEV-001. Synthetic memory-gate measurements are classified as synthetic-proxy and remain separate from the first-party observed live-model results. This dataset is governed by GC-MRD-v2.0 and evaluates the benchmark software archived at https://doi.org/10.5281/zenodo.21969841. The corresponding immutable GitHub release is available at https://github.com/RobbieRazor/robbies-razor-benchmarks/releases/tag/benchmark-results-v0.1.0. Evidence boundary: This record does not establish independent empirical validation, universal validity, model certification, production certification, cross-domain validation, scientific consensus, or guaranteed reductions in tokens, latency, computational cost, or resource use.

提供机构:
Zenodo
创建时间:
2026-08-17
二维码
社区交流群
二维码
科研交流群
商业服务