遇见数据集

Wild-OMR: A Public Phone-Camera Benchmark for Optical Mark Recognition Under Controlled Capture Variation

收藏
Zenodo2026-07-30 更新2026-08-02 收录
官方服务:

资源简介:

Wild-OMR is a public benchmark for optical mark recognition (OMR) onphotographs of answer sheets, built to measure where reading accuracy is lostwhen a bubble sheet is photographed with a phone instead of scanned. CONTENTS769 released captures of 30 pencil-filled answer sheets, each sheet carrying 40five-option questions (200 printed bubbles). Five phones spanning seven modelyears and three price tiers were used (Samsung Galaxy S26 Ultra, S24 Ultra andS24 FE; Xiaomi Redmi 13 and Redmi Note 8 Pro), plus one flatbed reference passat 300 dpi over the same physical sheets. The archive also ships the parametricsheet generator, the capture protocol, a synthetic generator with matchedrenderings of the same sheets, five classical reader configurations, threelearned baselines with their training scripts, per-question predictions forevery configuration, and the evaluation harness that produced every number inthe accompanying paper. DESIGNCapture varies one factor at a time rather than factorially: each conditiondeparts from a flat reference pose along a single axis. The five posedconditions are flat, dim, glare, tilted (about 30 degrees) and hand-held (sheetheld in the hand and allowed to curve). Illumination is therefore varied onlyon flat paper and geometry only under normal light, so no interaction betweenthem can be estimated from this data. Full design: 5 phones x 5 conditions x 30sheets + 30 flatbed frames = 780 captures. GROUND TRUTH AND PRIVACYGround truth is exact rather than annotated. Volunteers transcribed apre-assigned synthetic answer key onto anonymous, serial-numbered sheets, sothe released key file IS the per-bubble ground truth, with no annotation stepand no annotation error. Because nobody answered questions of their own, thesheets record a transcription rather than a performance: no names, no studentresponses and no personal data of any kind were collected. Twenty sheets arefilled cleanly; five carry a scripted erasure (a recorded decoy option filled,erased, then the intended answer filled) and five carry faint marks at roughlyhalf normal darkness, all annotated per question. PROVENANCESheet identity always comes from the sheet itself. Every released frame decodedits printed QR identifier. Of the 780 frames captured, 11 did not decode andwere dropped rather than relabelled: recovering their identity from captureorder was measured and would have worked, but an identity resting on aninference rather than on the sheet would weaken the one guarantee thisbenchmark offers. The dropped frames fall in the dim (4), tilted (6) andhand-held (1) conditions. LICENSINGImages, annotations, metadata and predictions are released under CC BY 4.0;the code is released under the MIT licence. See LICENSE-DATA and LICENSE-CODEinside the archive.

提供机构:
Zenodo
创建时间:
2026-07-30
二维码
社区交流群
二维码
科研交流群
商业服务