Reachable but not recognized: data and code for a two-vantage census and recognition study of 63 live religious radio streams, August 2026
收藏资源简介:
Companion data and code deposit of the article "Reachable but not recognized. Cultural coverage limits of general-purpose audio models on live religious radio" (Informatics Studies, special issue on Heritage Informatics and Religious Cultural Heritage, submitted September 2026). Between August 14 and August 24, 2026 a frozen catalog of 63 publicly advertised live religious radio streams across ten traditions, 21 languages and 21 countries was observed from two vantage points (a residential connection in Italy and GitHub Actions cloud runners in North American Azure regions) and sampled eight times a day, for 81 sessions and 4,570 analyzed samples of nominally 60 seconds (48 captures are shorter, between 1.1 and 51.0 s) from the 59 stations that delivered audio, with no pipeline failures. The cloud availability probe continued until September 20, 2026. The deposit contains the frozen station catalog, the two domain vocabularies written for CLAP zero-shot attribution with their cached prompt embeddings, the campaign and analysis scripts, every derived table (availability census and failure taxonomy, acoustic descriptors, the three-condition attribution experiment with out-of-sample split and stratification by acoustic form, paired McNemar tests, liturgical calendar contrasts, the longitudinal cloud series), the raw availability series of both vantage points, the captures manifest, the human listening sheet used to validate the acoustic-form indicator, the record of the frozen instrumentation with checkpoint fingerprints, and one derived analysis record per sample (technical descriptors, ecoacoustic indices, top AudioSet classes from PANNs CNN14, CLAP tags and the CLAP audio embedding as float16), which allows re-scoring the whole corpus against new text vocabularies without any audio. Audio recordings are not redistributed (copyrighted broadcast material) and speech transcripts were removed. Analysis toolkit: soundscape-audio-analysis v0.19.3 (tag v0.19.3), documented in https://doi.org/10.5281/zenodo.20282495. Every number, table and figure of the article regenerates from these files; see README.md.



