Visual-Taxonomic Informativeness (VTI): Supplementary Material, Reproducibility Package, and Software Reference Implementation
收藏资源简介:
Visual-Taxonomic Informativeness (VTI) is an expert-informed framework for assessing recoverable visual-taxonomic evidence in citizen-science biodiversity images and supporting repository auditing, evaluation-set construction, and target-aware training-data curation. The package contains the technical Supplementary Material, the experiment-specific reproducibility package used to obtain and validate the reported results, and a documented software reference implementation of the VTI framework. The reproducibility materials include experimental manifests, model registries, expert-informed visual ambiguity clusters, frozen ONNX backbone models, precomputed embeddings, analysis scripts, trained artifacts, result tables, bootstrap outputs, and other derived results supporting inspection and reproduction of the computational workflow. The study uses 120 fungal species and 120,000 citizen-science images, with DenseNet121 and EfficientNet-Lite0 retained as the principal frozen visual representations. The software reference implementation provides two user-facing workflows: offline calibration and operational scoring. Offline calibration constructs reference VTI targets from labelled images and fixed-probe predictions, optionally incorporating a user-provided expert-informed ambiguity registry, and then trains an embedding-based VTI predictor. Operational scoring applies the resulting calibration to additional images and produces image-level predicted VTI values without requiring species labels or fixed-probe outputs for the candidate images. The distributed archive is organized into three top-level components: Supplementary_Material.pdf, reproducibility_package/, and software_reference_implementation/. The software implementation includes documentation, dependency specifications, frozen ONNX models, model metadata, and citation information. Raw iNaturalist images are not redistributed. They remain subject to the licensing, availability, and reuse conditions of the source platform and individual content providers. Experimental manifests retain the identifiers and metadata required to relate the distributed computational artifacts to the source data. Precomputed embeddings used by the study are included in the Zenodo package and can also be regenerated from independently obtained source images using the supplied frozen models and scripts. The software reference implementation is domain-independent and does not include a preconfigured application domain or expert ambiguity registry. Calibration inputs are supplied by the user.



