Ontology curation of plant NGS metadata
收藏资源简介:
Run-level ontology annotations for plant next-generation sequencing records in the European Nucleotide Archive, covering the Viridiplantae clade in the snapshot synchronised through 16 April 2026. One table, dataset.tsv.gz, with one row per annotated run: 2,440,447 runs carrying 6,039,845 annotations on four axes. Tissue and developmental stage come from the Plant Ontology, treatment condition from the Plant Experimental Conditions Ontology, and genotype class from a small EFO and BAO vocabulary. The axes hold 32, 20, 42 and 8 distinct terms and annotate 1,863,364, 1,012,301, 1,573,429 and 1,590,751 runs. Each axis has four columns: the term, its identifier, how it was assigned and a confidence score. Where a run has no term on an axis, those four cells are empty. Each axis holds at most one term per run. On the treatment and genotype axes, terms whose source begins with default_ were filled in where no positive evidence was found. Positive evidence backs every tissue and developmental-stage term, 117,539 treatment terms and 198,226 genotype terms. These are automated assignments and some are wrong. README.md gives what each column means and what the file can and cannot be used to count. ENA metadata is not redistributed. Every row carries its ENA run accession, so the annotations can be matched back to ENA.



