The Sea Trial: A Maritime Honesty Benchmark — Abstention-Precision and Citation-Validity over Primary Instruments (v0.1.0)
收藏资源简介:
A benchmark that scores HONESTY, not fluency, in maritime AI. 284 items in three classes: CITED (170 — verifiable against named primary-source clauses: 33 CFR MTSA/ISPS regime, BIMCO Laytime Definitions 2013), NULL-TRAP (53 — the only correct output is an explicit abstention: fabricated premises, commercially-enclosed data, out-of-domain), COMPUTE (61 — deterministic engine ground truth with inputs echoed: laytime/demurrage, COLREG give-way, IMDG segregation, DG declaration completeness, grain stability, D&D tariffs, freight quotes). Score = accuracy × citation-validity × abstention-precision; a confident wrong answer scores below an abstention on every item. The scorer is fully deterministic (no LLM in the scoring path), calibrated by fixtures (oracle 1.0 / abstain-everything 0.25 / confident-fabricator 0.0). To the best of a dated search (2026-07-16), this is the first maritime benchmark scoring honesty rather than knowledge accuracy alone; maritime knowledge benchmarks (MaritimeBench) and general-domain abstention benchmarks (AbstentionBench) exist and are cited as validation of the premise. The deposit includes ANKR's own grounded-substrate self-score (final 0.1381) WITH the gap ledger it surfaced — the publisher publishes only its own scores, measured against its own prior release; third parties self-run the open harness. Harness code AGPL-3.0; item sets and documentation CC-BY-4.0.



