A deterministically linked corpus of U.S. Coast Guard marine casualty investigations, public AIS vessel tracks, and vessel regulatory history, 2024-2025
收藏资源简介:
A fully deterministic, auditable corpus linking three public U.S. sources across two near-complete calendar years, 2024 and 2025: closed U.S. Coast Guard marine casualty investigations (CGMIX Incident Investigation Reports), public MarineCadastre Automatic Identification System (AIS) vessel tracks, and per-vessel regulatory history from the Coast Guard's Port State Information Exchange (PSIX). No model, no language model, and no human adjudication enters the label path. Incident-to-track matching uses exact normalised identifier and vessel-name rules under fixed distance gates; every candidate is promoted to a confirmed case only by fixed multi-identifier corroboration (flag/MID and vessel class/AIS type) that resolves vessel-name collisions in code. The release contains 2,680 corroborated casualty cases drawn from 2,542 investigations and 2,046 distinct hulls, scanned from ~5.57 billion AIS messages over 730 days of nationwide archive coverage, together with 31,884 dated PSIX regulatory activity records across 1,721 vessels, resolved under one rule that is identical for casualty and control vessels and dated so they can be joined as of any score time. Every candidate, accepted or rejected alike, is published with its full matching evidence and disposition, so reusers can reconstruct the recommended casualty set or define their own label policy. Source works (USCG CGMIX; NOAA/BOEM/USCG MarineCadastre) are U.S. Government public domain; derived files are released under CC0-1.0. This record holds the corpus and the evidence behind it, together with a snapshot of the pipeline that produced it, released separately under the MIT license and included under code/. Results computed from the corpus are deliberately not included, so that reusing the data never means downloading someone else's findings. How to download. The corpus is deposited as nine tarballs so that most reusers do not have to fetch all of it. Start with the core part, which is self-contained for everything except the normal-traffic pool: the casualty records, the full candidate set with matching evidence and dispositions, vessel particulars, per-vessel-day summaries, incident-vessel AIS tracks, reference geometry, validation evidence, and the pipeline code. The eight normal-pool parts (one per calendar quarter) hold the deterministic normal-traffic sample and are needed only to reproduce results that use the negatives. Every part unpacks into the same directory, so extracting them together in one place reconstitutes the complete tree; PARTS.json in the core part maps paths to parts, and MANIFEST.json lists sha256 and length for every file across all parts.



