Supporting Information for "Forty New Genomes Shed Light on Sexual Reproduction and the Origin of Tetraploidy in Microsporidia"
收藏资源简介:
This is a list of the files and materials contained in this Supporting Information dataset. Supporting Information Section 1 Table S1: Microsporidian genome assemblies. Full list of recovered microsporidian genome assemblies and their associated meta-data. Genome accessions will be added as genomes are released through ENA. Supporting Information Table S1 File Collection S1: Microsporidian genome assemblies fasta files. Fasta files of recovered microsporidian genome assemblies. The primary assemblies listed in Table S1 are given by {Host ToLID}.µ.fasta, whereas purged haplotypic duplication sequences are given by {Host ToLID}.µ.alt.fasta where applicable. In the case of iuLoeVari1.µ, the primary assembly is the best haploid representative genome assembly possible, containing sequences across all four compartments. The diploid genome assemblies of iuLoeVari1.µ’s AB and CD compartments are given by iuLoeVari1.µ.AB.fasta and iuLoeVari1.µ.CD.fasta respectively. Supporting Information File Collection S1 Table S2: Filtering parameters used in generating genome assemblies. Parameters used for filtering microsporidian contigs from their respective (meta-)genomic assemblies in filtering steps 1 (BlobToolKit (Challis et al. 2020) and 2 (BubblePlot, Github: https://github.com/Amjad-Khalaf/BubblePlot). See Materials and Methods for details. Supporting Information Table S2 File Collection S2: Statistics of intermediate steps for each microsporidian genome assembly. Scaffold/contig and read statistics for intermediate steps produced in the generation of each microsporidian genome assembly. Supporting Information File Collection S2 File Collection S3: K-mer analysis plots for the microsporidian genome assemblies. MerqurkyFK plots (Github: https://github.com/thegenemyers/MERQURY.FK) for final microsporidian genome assemblies generated in this study. Supporting Information File Collection S3 File Collection S4: K-mer histogram plots for the reads used to produce the microsporidian genome assemblies. GenomeScope2 (Ranallo-Benavidez et al. 2020) plots of the reads used to produce the microsporidian genome assemblies presented in this study. Jellyfish was used to generate the k-mer spectrum for each read set (k = 21, version 2.2.10) (Marcais and Kingsford 2011). Supporting Information File Collection S4 File Collection S5: Smudgeplot ploidy estimation. Smudgeplot (Ranallo-Benavidez et al. 2020) plots of the reads used to produce the microsporidian genome assemblies for which ploidy could be estimated using GenomeScope2 (Ranallo-Benavidez et al. 2020) (Supporting Information File Collection S4). Supporting Information File Collection S5 File Collection S6: K-mer plots used to inform genome assembly purging. Purge_dups (Guan et al. 2020) histogram plots used to inform genome assembly purging, with cutoffs used clearly indicated. Supporting Information File Collection S6 File Collection S7: Self alignment plots for microsporidian genome assemblies. Self-alignment dot plots for the microsporidian genome assemblies generated in this study. Each genome was aligned to itself using FASTGA (Github: https://github.com/thegenemyers/FASTGA). The plots were generated using HyraxDotPlot (v2.0) (Github: https://github.com/Amjad-Khalaf/HyraxDotPlot). Supporting Information File Collection S7 File Collection S8: Hi-C contact heatmaps for scaffolded microsporidian genome assemblies. Hi-C contact heatmaps for scaffolded microsporidian genome assemblies visualised using PretextView (Harry 2022). Supporting Information File Collection S8 File Collection S9: Oxford Dot Plots for tetraploid microsporidian genome assemblies. Oxford dot plot for tetraploid genome assemblies displaying BUSCO genes. Gene pairs which are less divergent than the same species threshold are in sky blue, while gene pairs which are more divergent than the same species threshold are in red. Supporting Information File Collection S9 File Collection S10: Annotation results for microsporidian genome assemblies. Results for gene annotation (with Prokka) and repeat annotation (with RepeatModeler and RepeatMasker) on all genomes (Flynn et al. 2020; Seemann 2014; Smit et al. 2013). Supporting Information File Collection S10 Table S3: Accession numbers for publicly available genomes used in this study. On the 1st of January 2025, we downloaded all microsporidian genome assemblies available in the NCBI Genome database. This retrieved 106 genome assemblies. Supporting Information Table S3 Supporting Information Section 2 Table S4: Trait-phylogeny regression. Transformations representing the fit with the tree’s topology (λ), branch-lengths (κ) and root-tip distance (δ) (Pagel 1994) and the number of coding sequences, transposable element loads, and genome spans. Supporting Information Table S4 Table S5: Trait correlation. Correlations between the number of coding sequences, transposable element loads, and genome spans. Supporting Information Table S5 Supporting Information Section 3 Table S6: Branch length distances for species delineation. Pairwise branch length distances which include one of our genomes, and can be classified to a species or a genus. The conservative branch length threshold range was defined using the shortest observed branch lengths between known same-species genomes for the lower bound (0) and the smallest distance between H. tvaerminnensis and H. magnivora genomes for the upper bound (0.012). The relaxed threshold uses the full range of observed branch lengths among known same-species genomes (excluding the H. tvaerminnensis – H. magnivora cutoff). Supporting Information Table S6 Fig. S1: Comparison of whole-genome phylogeny species delineation thresholds and individual gene phylogeny branch length distribution species delineation thresholds when looking at tetraploid subgenomes. The approach we presented in Fig. 3 relies on branch lengths derived from the whole-genome phylogeny in Fig. 2 (i.e. a concatenated supermatrix of genes). We re-estimated same-species branch length thresholds for each gene. For each gene, we used the distribution of branch lengths between genomes known to belong to the same species, and measured each distribution’s mean and 95th percentile. The upper threshold was then set by retrieving the highest observed 95th percentile (orange dashed line), and the highest observed mean (magenta dashed line). While the percentage of genes exceeding each threshold varies for each genome, they are relatively consistent, and lead to the same OTU assignment and the same conclusions when investigating tetraploid species. ilAceEphe1 still stands out as possessing more genes which exceed the same-species threshold (no matter what threshold was used) than other genomes. Supporting Information Figure S1 Fig. S2: Relationship between whole-genome phylogeny species delineation thresholds and individual gene phylogeny branch length distribution species delineation thresholds. We compared our two gene-based metrics (highest 95th percentile and highest mean of branch length distributions of individual gene trees for genomes known to belong to the same species) to the whole-genome-based metric (highest branch length observed between any two same species genomes). We found the relationship between them to be consistent and linear, in line with the fact that they lead to the same conclusions. Supporting Information Figure S2 Supporting Information Section 4 Fig. S3: Tetraploid ilAceEphe1.µ is uneven and rearranged. The number of BUSCO genes found in X haplotypes, along with their total copy number. idChiSpeb1.µ is an even tetraploid, so nearly all its BUSCO genes are in 4 copies, distributed across 4 haplotypes. On the other hand, ilAceEphe1.µ is an uneven tetraploid. The majority of its BUSCO genes are in less than 4 copies, and they are not evenly distributed across its haplotypes. For instance, some BUSCO genes occur in 3 copies present only in a single haplotype. Supporting Information Figure S3 Supporting Information Section 5 Fig. S4: Phylogeny used by Syngraph, with its internal node labelling. Each node is labelled with its Syngraph name in a grey box. Yellow boxes indicate the number of chromosomes each genome possesses, and blue boxes indicate the number of chromosomes which possess BUSCO gene markers. Supporting Information Figure S4 Fig. S5: Number of chromosomes inferred at each node is highly variable. The number of chromosomes inferred for each node, and the total number of BUSCO genes assigned to a chromosome for each “m”. “m” is the parameter in Syngraph to determine the minimum number of genes needed to travel together for the event to be counted as a rearrangement. For example, if m = 3, only rearrangements involving 3 or more genes will be counted. Deep nodes are highly variable and their karyotype (and thus the number of rearrangements that have occurred along each branch) cannot be estimated reliably. See Fig. S4 for node labels on the phylogeny. Supporting Information Figure S5 Fig. S6: T-SNE plot depicting BUSCO linkage groups across the microsporidian phylogeny. Each point represents a BUSCO gene, positioned based on its co-occurrence profile across the chromosome-level microsporidian genomes. Distances between points reflect similarities in co-occurrence. Points are coloured by their assigned chromosome in Anotonspora locustae. This disorganised pattern illustrates that the rate of rearrangement is too high for a reliable complete reconstruction of putative ancestral linkage groups. The large-scale patterns are influenced by more densely sampled taxa, see Fig. S7. Supporting Information Figure S6 Fig. S7: T-SNE plot depicting BUSCO linkage groups across the microsporidian phylogeny, highlighting clustering influence by more densely sampled taxa. Each point represents a BUSCO gene, positioned based on its co-occurrence profile across the chromosome-level microsporidian genomes. Distances between points reflect similarities in co-occurrence. Points are coloured by their assigned chromosome in Encephalitozoon cuniculi. This disorganised pattern illustrates that the rate of rearrangement is too high for a reliable complete reconstruction of putative ancestral linkage groups. The large-scale patterns are influenced by more densely sampled taxa, such as Encephalitozoon cuniculi. Supporting Information Figure S7 Fig. S8: Synteny plots of chromosomal microsporidian genome assemblies. Genome-wide synteny plots of all available chromosomal microsporidian genome assemblies. Each line represents a single-copy BUSCO (microsporidia_odb10) (Simao et al. 2015). BUSCOs are painted by their chromosomal position in A. locustae. Plot was generated by modifying ribbon plot scripts from https://github.com/conchoecia/odp (Schultz et al. 2023). Supporting Information Figure S8 Fig. S9: Synteny plots of chromosomal microsporidian genome assemblies. Genome-wide synteny plots of all available chromosomal microsporidian genome assemblies. Each line represents a single-copy BUSCO (microsporidia_odb10) (Simao et al. 2015). BUSCOs are painted by their chromosomal position in H. tvaerminnensis. Plot was generated by modifying ribbon plot scripts from https://github.com/conchoecia/odp (Schultz et al. 2023). Supporting Information Figure S9



