遇见数据集

VARF: Synthetic Audiogram Generator, Experiment Code and Systematic Review Materials

收藏
Zenodo2026-08-09 更新2026-08-13 收录
官方服务:

资源简介:

Supporting materials for the manuscript "VARF: A Five-Layer AI Architecture forVirtual Aural Rehabilitation: Design and Algorithmic Feasibility", submitted toEngineering, Technology & Applied Science Research. Contents. A deterministic synthetic audiogram generator (seed 20260301) and thecomplete code for four experiments: a hybrid rule-plus-gradient-boosted triageclassifier, adversarial robustness checks, differentially private federatedaveraging with Renyi differential-privacy accounting, and a LinUCB adaptivedifficulty controller benchmarked against a clinical staircase procedure. Alsoincluded: the generated dataset of 5,000 cases, a reward-weight sensitivityanalysis, a privacy-accountant configuration scan, the systematic reviewmaterials (database-specific search strings for PubMed, Scopus, CINAHL and IEEEXplore with all deviations recorded, the PRISMA 2020 flow, and the coding rubricunderlying Table II), and the six figures as published. Provenance. varf_reproduce.py is the authoritative specification of the method:every numerical value in Section III of the manuscript is its output.PROVENANCE.md records which reported values changed relative to the originallysubmitted version, and why. Important caveats. All results derive from synthetic data. No record-levelpatient data and no record-level NHANES data were used at any stage; thegenerator's parameters were calibrated to published summary statistics only,namely age- and frequency-stratified mean thresholds and standard deviations fromthe NHANES audiometry examination and the WHO hearing-impairment grade cut-offs.Because the triage labels and the inference-time rule set derive from the sameclinical criteria that parameterised the generator, the reported metrics indexrecoverability of a deliberately encoded structure and are an upper bound on whatthe pipeline could achieve on real audiograms, not an estimate of it. No clinicalvalidity is claimed. One result is negative. The LinUCB difficulty controller is outperformed by atransformed two-down-one-up staircase, the adaptive procedure clinicalspeech-in-noise testing already uses (0.767 against 0.703 mean reward; 90.6%against 52.1% of trials within the target zone). The manuscript reports this as anegative result and frames the corresponding architectural layer as an interfacespecification rather than as evidence for the bandit. Headline results. Triage classifier 0.938 +/- 0.008 accuracy, 0.971 macro-AUC,expected calibration error 0.038 under 5-fold cross-validation. Differentiallyprivate federated learning 91.0% accuracy at epsilon approximately 4.23 and 88.1%at epsilon approximately 1.44, against a non-private centralised baseline of93.5%. Licences. Data, figures and documentation under CC BY 4.0; code under MIT. SeeLICENSE. Generative AI. The code was written with the assistance of a generative AI tool(Claude, Anthropic). This constitutes use in the research process rather thanlanguage editing and is disclosed as such in the manuscript's AI declaration. Reproducing. Download all files into one directory and run ./run_all.sh, aboutfour minutes on a commodity CPU. Output is deterministic and has been verifiedidentical on two independent machines: compare varf_out/metrics.json againstmetrics_reference_run.json. Version 1.2.0 note. This version withdraws the four per-study review templatefiles carried by earlier versions. Extraction and quality-assessment records werenot retained in a form suitable for deposit, and reconstructing them after thefact would produce material indistinguishable from a contemporaneous record. Thereview is therefore documented at the level of its protocol, search strategy,screening counts and exclusion reasons, in SEARCH_STRINGS.md, PRISMA_FLOW.md andCODING_RUBRIC.md, and not study by study; the modal codes in Table II of themanuscript cannot be recomputed independently from this archive. The manuscriptstates this as a limitation. It attaches to the review component alone: thegenerator, the four experiments and every value in Section III reproduce in fulland deterministically from the code deposited here.

提供机构:
Zenodo
创建时间:
2026-08-09
二维码
社区交流群
二维码
科研交流群
商业服务