OIP-Bench v0.2.0: frozen model replies and benchmark packages of the Open Image Protocol paper
收藏资源简介:
Frozen replies of eleven vision-language models (three frontier APIs, six Ollama cloud models, two local models) to the OIP-Bench understanding tasks on 100 VinDr-CXR radiographs, 50 NIH ChestX-ray14 radiographs and 40 whole-body bone scans, collected 8–22 September 2026: 41 run directories with results.jsonl and tasks.json (67,092 rows), plus one run excluded under the no-leakage rule. Benchmark packages: full OIP packages for NIH ChestX-ray14 and the Paraguay bone scans; for VinDr-CXR (Kaggle competition data, not redistributable) the manifests, reference files and measurements only, with sha256 hashes of every file so a rebuild can be verified. MANIFEST.json lists every bundled file with sha256 and size. Rows carry the scorer-0.3 result in score and the scorer-0.2 result used in the paper in score_prev. How to reproduce the paper's tables from these files: REPRODUCE.md in the repository, level B.



