CAVEAT: Collaborative Aided Target Recognition with Evidence-Grounded Verification and Adjudication for Multi-Domain Unmanned Systems
收藏资源简介:
Heterogeneous unmanned teams now operate together, but each platform's perception operates alone. So, it is necessary for UAS, UGV, and USV to detect, identify, and disambiguate targets at long range under appearance ambiguity, viewpoint and scale variation, sensor heterogeneity across EO/IR, SAR, LiDAR, and RF, and degraded position, navigation, and timing (PNT). Reconciling what several platforms saw is today a largely manual process: slow, communication-intensive, and error-prone exactly when error is least affordable. Computer vision Foundation models make collaborative aided target recognition technically plausible now with vision-language models (VLM). VLM models make open-vocabulary predictions and make it possible to identify over long, asynchronous multi-platform streams tractable on edge devices. But VLM models are generally pre-trained on web-scale civilian data sets rather than the military target set. They often struggle to generalize reliably in operational settings and may generate highly confident but incorrect predictions. For example, a vision-language model can see a bird and predict a quadcopter drone at 0.95 confidence, citing “two propellers” that do not exist. This is a NEUTRAL declared as a THREAT.



