遇见数据集

V2 Supporting Information for "Forty New Genomes Shed Light on Sexual Reproduction and the Origin of Tetraploidy in Microsporidia"

收藏
Zenodo2025-08-05 更新2026-05-26 收录
官方服务:

资源简介:

This is a list of the files and materials contained in this Supporting Information dataset (v2). Section 1 Fig. S1: Sex of hosts the microsporidian genome assemblies are derived from. The sex of our genomes’ hosts was unknown in most cases (24 species). In the remaining cases, nine were identified as female and seven as male. A relatively equal proportion of female and male hosts are infected with Nosematida (Fig. 1), but we could not assess skews in host sex ratios for other microsporidian groups due to missing data on sex (for Amblyosporida-infected hosts), or a small sample size (for Neopereziida-infected hosts). Table S1: Microsporidian genome assemblies. Full list of recovered microsporidian genome assemblies and their associated meta-data. Genome accessions will be added as genomes are released through ENA. File Collection S1: Microsporidian genome assemblies fasta files. Fasta files of recovered microsporidian genome assemblies. The primary assemblies listed in Table S1 are given by {Host ToLID}.µ.fasta, whereas purged haplotypic duplication sequences are given by {Host ToLID}.µ.alt.fasta where applicable. In the case of iuLoeVari1.µ, the primary assembly is the best haploid representative genome assembly possible, containing sequences across all four compartments. The diploid genome assemblies of iuLoeVari1.µ’s AB and CD compartments are given by iuLoeVari1.µ.AB.fasta and iuLoeVari1.µ.CD.fasta respectively. Table S2: Filtering parameters used in generating genome assemblies. Parameters used for filtering microsporidian contigs from their respective (meta-)genomic assemblies in filtering steps 1 (BlobToolKit [124]) and 2 (BubblePlot, Github: https://github.com/Amjad-Khalaf/BubblePlot). See Materials and Methods for details. File Collection S2: Statistics of intermediate steps for each microsporidian genome assembly. Scaffold/contig and read statistics for intermediate steps produced in the generation of each microsporidian genome assembly. File Collection S3: K-mer analysis plots for the microsporidian genome assemblies. MerqurkyFK plots (Github: https://github.com/thegenemyers/MERQURY.FK) for final microsporidian genome assemblies generated in this study. File Collection S4: K-mer histogram plots for the reads used to produce the microsporidian genome assemblies, used for ploidy estimation. GenomeScope2 [81] plots of the reads used to produce the microsporidian genome assemblies presented in this study. Jellyfish was used to generate the k-mer spectrum for each read set (k = 21, version 2.2.10) [126]. File Collection S5: Smudgeplot ploidy estimation. Smudgeplot [81] plots of the reads used to produce the microsporidian genome assemblies for which ploidy could be estimated using GenomeScope2 [81] (Supporting Information Section 1, File Collection S4). Fig. S2: Ploidy inference examples for three microsporidian genomes, highlighting segmental duplications. GenomeScope2 transformed linear plot and Smudgeplot [81] respectively for (A), (B) diploid iyCepSpine2.µ (host Cephus spinipes [Hymenoptera]); (C), (D) diploid idPhaFune2.µ (host Phania funesta [Diptera]); and (E), (F) polyploid (tetraploid or octoploid) iiMysAzur1.µ (host Mystacides azureus [Trichoptera]). Jellyfish was used to generate the initial k-mer spectra (k = 21, version 2.2.10) [126]. Both iyCepSpin2.µ and idPhaFune2.µ have mostly diploid genomes, but carry a level of duplication that generated an identifiable “tetraploid” signal in their k-mer spectra. Similarly, the k-mer spectrum of iiMysAzur1.µ can be interpreted as either a highly homozygous tetraploid where large segmental duplications have occurred in all the four copies leading to a detectable octoploid signal, or an octoploid genome composed of two distinct tetraploids. Such cases are common, with some level of segmental duplication observed in nearly all of the 14 polyploid genomes (refer to Supporting Information Section 1, File Collection S4). File Collection S6: K-mer plots used to inform genome assembly purging. Purge_dups [74] histogram plots used to inform genome assembly purging, with cutoffs used clearly indicated. File Collection S7: Self alignment plots for microsporidian genome assemblies. Self-alignment dot plots for the microsporidian genome assemblies generated in this study. Each genome was aligned to itself using FASTGA (Github: https://github.com/thegenemyers/FASTGA). The plots were generated using HyraxDotPlot (v2.0) (Github: https://github.com/Amjad-Khalaf/HyraxDotPlot). File Collection S8: Hi-C contact heatmaps for scaffolded microsporidian genome assemblies. Hi-C contact heatmaps for scaffolded microsporidian genome assemblies visualised using PretextView [94]. File Collection S9: Oxford Dot Plots for tetraploid microsporidian genome assemblies. Oxford dot plot for tetraploid genome assemblies displaying BUSCO genes. Gene pairs which are less divergent than the same species threshold are in sky blue, while gene pairs which are more divergent than the same species threshold are in red. File Collection S10: Annotation results for microsporidian genome assemblies. Results for repeat annotation (with RepeatModeler and RepeatMasker) on all genomes [79,80]. Table S3: Accession numbers for publicly available genomes used in this study. On the 1st of January 2025, we downloaded all microsporidian genome assemblies available in the NCBI Genome database. This retrieved 106 genome assemblies. Section 2 Table S4: Trait-phylogeny regression. Transformations representing the fit with the tree’s topology (λ), branch-lengths (κ) and root-tip distance (δ) [82] and the number of coding sequences, transposable element loads, and genome spans. Table S5: Trait correlation. Correlations between transposable element loads, and genome spans. Text S1: Newick string of phylogeny. ASTRAL [76] phylogeny summarising individual phylogenies of 600 BUSCO genes (microsporidia_odb10) [77] across all publicly available microsporidian genome assemblies (n = 106), and the genome assemblies generated in this study (n = 40, marked in purple). Branch lengths were estimated with IQ-TREE using a concatenated alignment of the individual BUSCOs [78]. The model chosen according to IQ-TREE’s model finder was “Q.yeast.I.G4”. Fig. S3: 600 Gene Phylogeny of Microsporidia. (A) ASTRAL [76] phylogeny summarising individual phylogenies of 600 BUSCO genes (microsporidia_odb10) [77] across all publicly available microsporidian genome assemblies (including multiple strains where they are available), and the genome assemblies generated in this study (n = 40, marked in purple). Branch lengths were estimated with IQ-TREE using a concatenated alignment of the individual BUSCOs [78]. Nodes with less than 95% support are marked with pink circles. Ploidy is marked in circles at the tips of the tree for genomes where it was characterisable. (B) Genome assembly span (Mb), with black circles marking chromosome-level genome assemblies. (C) N50 values (Mb), with asterisks marking purged genome assemblies. (D) BUSCO gene (microsporidia_odb10) completeness percentage, marked in green for single-copy genes, and beige for duplicated genes. (E) Transposable element percentage as predicted by RepeatModeler and RepeatMasker [79,80], marked in burgundy for retroelements, peach for DNA transposons, and blue for rolling circles. Neop.: Neopereziida; Or. Lin.: Orphan Lineage. File Collection S11: Repeat landscape across chromosome-level microsporidian genome assemblies. Distribution of repeat annotation results (with RepeatModeler and RepeatMasker) on all chromosome-level genomes [79,80]. File Collection S12: rRNA landscape across chromosome-level microsporidian genome assemblies. Distribution of rRNA across all chromosome-level genomes. Fig. S4: Discrepancy between protein annotation approaches in Microsporidia. Each position on the x-axis represents a genome, with the y-axis showing annotation results. Different tools are represented by different colours. A wide discrepancy is observed between different tools, and between hypothetical annotation approaches and the annotations retrieved from Genbank [132–135]. Section 3 Table S6: Branch length distances for species delineation. Pairwise branch length distances which include one of our genomes, and can be classified to a species or a genus. The conservative branch length threshold range was defined using the shortest observed branch lengths between known same-species genomes for the lower bound (0) and the smallest distance between H. tvaerminnensis and H. magnivora genomes for the upper bound (0.012). The relaxed threshold uses the full range of observed branch lengths among known same-species genomes (excluding the H. tvaerminnensis – H. magnivora cutoff). Fig. S5: Comparison of whole-genome phylogeny species delineation thresholds and individual gene phylogeny branch length distribution species delineation thresholds. The approach we presented in the main text relies on branch lengths derived from the whole-genome phylogeny in Fig. 2 (i.e. a concatenated supermatrix of genes). We re-estimated same-species branch length thresholds for each gene. For each gene, we used the distribution of branch lengths between genomes known to belong to the same species, and measured each distribution’s mean and 95th percentile. The upper threshold was then set by retrieving the highest observed 95th percentile (orange dashed line), and the highest observed mean (magenta dashed line). While the percentage of genes exceeding each threshold varies for each genome, they are relatively consistent, and lead to the same OTU assignment and the same conclusions when investigating tetraploid species. ilAceEphe1.µ still stands out as possessing more genes which exceed the same-species threshold (no matter what threshold was used) than other genomes. Fig. S6: Relationship between whole-genome phylogeny species delineation thresholds and individual gene phylogeny branch length distribution species delineation thresholds. We compared our two gene-based metrics (highest 95th percentile and highest mean of branch length distributions of individual gene trees for genomes known to belong to the same species) to the whole-genome-based metric (highest branch length observed between any two same species genomes). We found the relationship between them to be consistent and linear, in line with the fact that they lead to the same conclusions. Section 4 Fig. S7: Tetraploid ilAceEphe1.µ is uneven and rearranged. The number of BUSCO genes found in X haplotypes, along with their total copy number. idChiSpeb1.µ is an even tetraploid, so nearly all its BUSCO genes are in 4 copies, distributed across 4 haplotypes. On the other hand, ilAceEphe1.µ is an uneven tetraploid. The majority of its BUSCO genes are in less than 4 copies, and they are not evenly distributed across its haplotypes. For instance, some BUSCO genes occur in 3 copies present only in a single haplotype. Section 5 Text S2: Details on rearrangements inferred and methods attempted. Comparing the BUSCO positions on closely-related genomes showed a pattern of dynamic change of microsporidian linkage groups. The Encephalitozoon species genomes were highly syntenic, with the exception of rearrangement involving a single linkage group [67] (Fig. S12-S13). ilEupExig1.µ (host Eupithecia exiguata [Lepidoptera]), placed basal to Encephalitozoon species, had six chromosomes that could be derived through either five pairwise fusions of the 11 chromosomes of Encephalitozoon species, or an ancestral karyotype that was subject to fission in Encephalitozoon. Vairimorpha necatrix was sister to the Encephalitozoon-ilEupExig1.µ clade, and had 11 chromosomes, but these did not correspond to the 11 found in Encephalitozoon species and did not simply confirm the karyotype of ilEupExig1.µ as being ancestral or derived. The karyotype of the enterocytozoonid iyOphElle1.µ, sister to the nosematids (V. necatrix, Encephalitozoon and ilEupExig1.µ), had 12 chromosomes, the largest of which is syntenic with the largest chromosome of V. necatrix, and thus suggests that the splitting of this chromosome in Encephalitozoon and ilEupExig1.µ is derived. This large linkage group was also found in the neoperezeiid Antonospora locustae, sister to the Encephalitozoonida, in the orphan lineage species Hamiltosporidium tvaerminnensis and in two genomes on Ambylosporidia species, suggesting it was likely present in the last ancestor of all microsporidia analysed. The orphan lineage-Ambylosporidia group showed a similar general pattern of linkage group conservation between close relatives and members of the same OTU, coupled with major rearrangements between clades (Fig. S12-S13). We were unable to infer a robust set of putative ancestral linkage groups for these genomes using syngraph [136] (Fig. S8-S9) or unsupervised clustering of loci based on their chromosomal occupancy [137,138] (Fig. S10-S11), likely because of the high frequency of rearrangements observed. Fig. S8: Phylogeny used by Syngraph, with its internal node labelling. Each node is labelled with its Syngraph name in a grey box. Yellow boxes indicate the number of chromosomes each genome possesses, and blue boxes indicate the number of chromosomes which possess BUSCO gene markers. Fig. S9: Number of chromosomes inferred at each node is highly variable. The number of chromosomes inferred for each node, and the total number of BUSCO genes assigned to a chromosome for each “m”. “m” is the parameter in Syngraph to determine the minimum number of genes needed to travel together for the event to be counted as a rearrangement. For example, if m = 3, only rearrangements involving 3 or more genes will be counted. Deep nodes are highly variable and their karyotype (and thus the number of rearrangements that have occurred along each branch) cannot be estimated reliably. See Fig. S8 for node labels on the phylogeny. Fig. S10: T-SNE plot depicting BUSCO linkage groups across the microsporidian phylogeny. Each point represents a BUSCO gene, positioned based on its co-occurrence profile across the chromosome-level microsporidian genomes. Distances between points reflect similarities in co-occurrence. Points are coloured by their assigned chromosome in Anotonspora locustae. This disorganised pattern illustrates that the rate of rearrangement is too high for a reliable complete reconstruction of putative ancestral linkage groups. The large-scale patterns are influenced by more densely sampled taxa, see Fig. S11. Fig. S11: T-SNE plot depicting BUSCO linkage groups across the microsporidian phylogeny, highlighting clustering influence by more densely sampled taxa. Each point represents a BUSCO gene, positioned based on its co-occurrence profile across the chromosome-level microsporidian genomes. Distances between points reflect similarities in co-occurrence. Points are coloured by their assigned chromosome in Encephalitozoon cuniculi. This disorganised pattern illustrates that the rate of rearrangement is too high for a reliable complete reconstruction of putative ancestral linkage groups. The large-scale patterns are influenced by more densely sampled taxa, such as Encephalitozoon cuniculi. Fig. S12: Synteny plots of chromosomal microsporidian genome assemblies. Genome-wide synteny plots of all available chromosomal microsporidian genome assemblies. Each line represents a single-copy BUSCO (microsporidia_odb10) [77]. BUSCOs are painted by their chromosomal position in A. locustae. Plot was generated by modifying ribbon plot scripts from https://github.com/conchoecia/odp [100]. Fig. S13: Synteny plots of chromosomal microsporidian genome assemblies. Genome-wide synteny plots of all available chromosomal microsporidian genome assemblies. Each line represents a single-copy BUSCO (microsporidia_odb10) [77]. BUSCOs are painted by their chromosomal position in H. tvaerminnensis. Plot was generated by modifying ribbon plot scripts from https://github.com/conchoecia/odp [100].

提供机构:
Zenodo
创建时间:
2025-08-05
二维码
社区交流群
二维码
科研交流群
商业服务