uavews: Dataset Formation, Technical Validation and Field-Trial Planning for Small-UAV Early Warning
收藏资源简介:
Description uavews is the versioned curation, validation and packaging pipeline for a multisource, multimodal spatiotemporal dataset on the early warning of approaching small unmanned aerial vehicles. It normalizes four heterogeneous source streams — takeoff indications from a controlled area, public warning events, voluntary mobile reports, and visual and acoustic records from authorized monitoring sites — into one event-centered model, and carries every record through ten stages: ingestion, window construction, uncertainty-expanded association, derivation of kinematic ground truth, conflict adjudication across three evidence tiers, de-identification and access tiering, technical validation, leakage-resistant partitioning, packaging, and application of the release gates. Each stage records a PROV-O activity, and every released file carries a SHA-256 digest. Reported quantities are measured rather than declared. Boundary distance and warning time are derived from the reference trajectory, with a direction dead-band computed from the positional uncertainty and enforced as a floor. Synchronization error is evaluated on the marker that opens each run, and sources that cannot observe a marker report their declared uncertainty and a null measurement rather than a fabricated one. Media quality is computed from the released bytes: the acoustic estimator is referred to the same in-band noise power the propagation model predicts, declares its own sensitivity bound, and returns a null instead of reporting noise as a weak detection. Duplicate rate is evaluated over perceptual object groups, and every evaluation manifest is audited against its own stated constraint rather than a generic one. The release in this record is a rehearsal, not a measurement campaign. The corpus it contains is synthetic, generated from a fixed seed, and carries deliberately injected defects — redelivered upstream events, re-encoded media, gross clock offsets, dropped channels, clipped and silent audio, incidental speech, annotator disagreement near the detection limit — so that each release gate is confirmed to fire on the condition it is meant to detect, before any hardware is deployed. No value it produces is a measurement, and none may be transcribed into the bracketed placeholders of the accompanying Data Descriptor. The empirical dataset will be deposited as a separate record and linked to this one. The archive contains the source code (23 modules, 54 tests), the parameter file and controlled vocabularies with which the release was produced, the raw corpus so that every published digest can be verified against actual bytes, the deposit-shaped RO-Crate package with DataCite and PROV-O metadata, a machine-readable validation report, 16 figures, and a twelve-section engineering report stating every calculation with a worked numeric example. The rehearsal release covers 180 events (90 controlled flights, 40 observational episodes and 50 hard negatives across five confounder families), 971 observations, 391 media objects and 7,242 released labels. Canonical tables and split manifests reproduce byte-for-byte across runs; the three files that differ all carry wall-clock timestamps. The package also sizes the campaigns that are to replace the rehearsal corpus, computing acoustic and visual detection ranges from declared physical assumptions, the operational warning-time budget those ranges buy, and the number of sorties the flight matrix requires. Reproduce with pip install -r code/requirements.txt and PYTHONPATH=code/src python -m uavews.cli all --out build.



