RE-SAPRIA Phase 5A: Standardized structural annotations of Sapria himalayana genome representations
收藏资源简介:
RE-SAPRIA Phase 5A provides standardized structural gene annotations for multiple genome representations of the endoparasitic flowering plant Sapria himalayana (Rafflesiaceae). Sapria himalayana has been independently sequenced in two published studies, represented here as the Cai and Guo datasets after the first authors of those studies. RE-SAPRIA investigates how much apparent genomic divergence between these representations reflects biological sequence variation versus genome reconstruction, repeat architecture, and annotation strategy. This deposit contains four RNA-supported BRAKER3 annotations generated from softmasked genome representations: Guo published genome; Cai fixed-on-Guo analytical pseudogenome; Cai published assembly filtered to scaffolds ≥1,250 bp; independently reconstructed Cai Flye–HyPo assembly. The same three Guo paired-end RNA-seq libraries were mapped independently to each genome representation and used as transcript evidence. No protein evidence was supplied to these standardized BRAKER3 runs. The deposit also contains two standalone AUGUSTUS control annotations for the Guo genome, generated using the same trained gene-prediction model while changing only repeat visibility: one softmasked and one unmasked. These files are intended as a methodological control for evaluating annotation sensitivity to repeat masking and should not be treated as the primary gene annotations. The Cai fixed-on-Guo sequence is an analytical pseudogenome in which confident Cai homozygous-alternate alleles were introduced onto the Guo structural backbone. It is not an independently assembled Cai genome. Phase 5A analysis shows that Guo and Cai fixed-on-Guo produce nearly identical global BRAKER annotation scales, while Cai min1250 and Cai Flye–HyPo retain similar gene and transcript counts despite substantial differences in assembly span. The standardized annotations also show that very long introns are a robust feature across genome representations. Repeat-aware interpretation of these annotations will be extended in RE-SAPRIA Phase 5B following completion of the corresponding RepeatMasker analyses. Files are provided as compressed GFF3 annotations together with SHA-256 checksums and a manifest describing their provenance and intended analytical role. This dataset is part of the ongoing RE-SAPRIA comparative-genomics project and should be interpreted as an active research dataset rather than as a finalized community reference annotation.



