A reuse-ready region map for Brazilian artisanal-cheese amplicon data: PCR primers, target regions and cutadapt actions for 287 sequencing runs across nine public BioProjects
收藏资源简介:
What this is. A curated resource that makes 287 publicly archived amplicon sequencing runs, from nine Brazilian artisanal-cheese BioProjects, actually reprocessable, together with the code that produced it. For each BioProject it gives the forward and reverse PCR primer sequences, the targeted region (16S V3-V4, V4, V3-V5; fungal ITS2), the library layout, the correct cutadapt behaviour, the document each primer sequence was read from, and a per-row confidence and verification status. Why it was needed. The archive names the marker gene or the target region for most studies, but it almost never records the primer itself. Measured across every cheese amplicon deposit in the ENA on 31 July 2026 (194 studies, 15,743 runs), no primer sequence appears anywhere in the record for 175 of 194 studies (90.21%; 95% Wilson CI 85.21-93.64), and only one study yields one through the tabular query interface. What the archive does record about the amplification is largely held in the <DESIGN_DESCRIPTION> element of the INSDC experiment schema, which the ENA Portal API does not expose: of the 123 studies whose record names the target, 120 (97.56%) name it only there. A reuser querying the tabular interface retrieves nothing and may conclude, wrongly, that nothing was recorded. Every primer pair in this map therefore had to be reconstructed from outside the archive's queryable fields. Contents (36 files). data/region_map.csv, one row per BioProject with primers, region, layout, cutadapt action, source document, source DOI, CrossRef verification status and confidence. code/, seventeen scripts, each annotated with the claim it regenerates: the archive-metadata tests, the whole-universe measurements, the detector and classifier validations, the universe-sensitivity tests, the FASTQ primer inspections, and the scripts that build the manuscript's table, figures and numeric audit. logs/, seventeen TSV files holding the raw output of those scripts. MANIFESTO.md gives the role, the supported manuscript section and the SHA-256 of every file. Each script takes the accessions as its only input and queries the ENA and NCBI directly, using no intermediate file from the project, so any reader can regenerate the claims independently. Six reuse hazards documented here and visible only in the raw FASTQ: primers already removed before deposit in three BioProjects (running cutadapt would discard essentially the whole study, or decapitate real bases); reads already merged by FLASH in two BioProjects despite a SINGLE layout label; a 0-5 nt heterogeneity spacer preceding the primer in one BioProject, so fixed-position trimming fails; a single degenerate base separating apparently identical V4 primer sets; two different sequences published under the same primer name, 341F, and one sequence published under two names, 785R and 805R; and a deposit whose reads carry a forward primer other than the one its linked publication declares. Structural limit, declared up front. After harmonisation, no region x primer x layout combination assembles more than two studies: the 287 runs fall into seven groups of 96, 72, 41, 35, 26, 13 and 4 runs, five of which contain a single study. Across the georeferenced subset, grid cell and BioProject are perfectly confounded (Cramer's V = 1.0000 at every scale from 1 to 100 km). This collection cannot support cross-study regional inference; it supports descriptive, harmonised reprocessing within each group and methodological comparison between groups. No new sequences were generated. All FASTQ files remain in ENA/SRA under the accessions listed. This deposit contains metadata, curation and code only. Verification. Every DOI in the region map and in the accompanying manuscript was re-checked against the CrossRef API (api.crossref.org/works/{DOI}) on 31 July 2026 against a first author, title fragment and year stated before the query: 15 of 15 resolve and match, none is unverified. Where CrossRef's issued date differed from the issue year, published-print was adopted. The proportions above are reported with explicit denominators and 95% Wilson confidence intervals, and a script in code/manuscript/ re-derives each one from the manuscript text and compares it against the arithmetic.



