Additional file 2 of A comprehensive atlas of nuclear sequences of mitochondrial origin (NUMT) inserted into the pig genome
收藏资源简介:
Additional file 2: Table S5. Information on the NUMT identified in the assembled genomes. Table S6. List of all NUMT sequences identified in the assembled genomes with LAST alignment information. Table S7. Comparison at the chromosome/scaffold level between NUMT identified in Sscrofa11.1 and Sscrofa10.2 reference genomes. Table S8. Matrix of orthologous NUMT regions over all Suinae investigated assembled nuclear genomes. The matrix below the diagonal contains the number of NUMT regions orthologous between the two genomes; the diagonal of the matrix reports the total number of regions found in the genome and the number of regions belonging to portions of the genome not aligned with Sscrofa11.1. The matrix above the diagonal contains the number of regions from genome A not found in genome B and vice versa. The acronyms of the reference genomes are explained in Additional file 2 Table S5. Table S9. List of all non-redundant NUMT regions found in the assembled genomes with presence/absence in all other assembled genomes. NUMT regions that could not be reconducted to a position in Sscrofa11.1 were not included. The first six columns contain the unique NUMT region ID and its coordinates in Sscrofa11.1. For NUMT regions absent in Sscrofa11.1 the coordinates of the insertion point are reported instead. The remaining columns refer to the information for the unique NUMT region in the assembled genomes: the reported information for each genome, when available, were the ID and coordinates of the corresponding NUMT region, along with its status, defined as follows: CORRESPONDING, if the genome contains the NUMT region; COMPATIBLE SEQUENCE, if the genome aligns with Sscrofa11.1 on the NUMT region coordinates, and the corresponding sequence in the other genome was not annotated as NUMT by the NUMT discovery pipeline; PRIVATE, if the alignment between the assembled genome and Sscrofa11.1 contains gaps corresponding to the NUMT region coordinates in one of the two genomes, meaning that the NUMT insertion happened only in one of the two genomes, and the NUMT can be therefore considered as polymorphic. Each status column contains the information for Sscrofa11.1 on the left side and the other genome on the right side. Table S10. List of all novel NUMT identified by mining WGS datasets. Breakpoints belonging to the same NUMT insertion are grouped with the same NUMT region ID. Table S11. Frequency of the carriers of novel NUMT regions estimated from WGS datasets within breed/population and species. For each NUMT region, number of WGS datasets and their percentageof the carriers is reported. Table S12. Genotype information obtained by PCR analyses of some polymorphic NUMT in several pig breeds and populations. The number of pigs having the three genotypesis reported. Table S13. Frequency of the carriers of all non-redundant NUMT regions found in assembled genomes estimated from WGS datasets within breed/populations and species. For each NUMT region, number of WGS datasets and their percentage of the carriers is reported. Table S14. Information on NUMT derived from the assembled genomes and WGS datasets with age estimation, presence/absence in the assembled genomes and WGS datasets divided by species and geographic origin of the pigs. Table S15. Comparative age estimation of NUMT regions assigned to the age classes 10–55 Mya and > 55 Mya, based on different mutation rates. Table S16. Annotation of the genomic regions where the NUMT regions are inserted.
附加文件2:表S5。组装基因组中鉴定到的核线粒体DNA片段(nuclear mitochondrial DNA segment, NUMT)的相关信息。 表S6。组装基因组中鉴定到的所有NUMT序列列表,附带LAST比对信息。 表S7。Sscrofa11.1与Sscrofa10.2参考基因组中鉴定到的NUMT在染色体/支架水平的比较分析。 表S8。所有已组装的猪亚科(Suinae)核基因组中直系同源NUMT区域的矩阵。矩阵对角线下方区域展示两个基因组间直系同源NUMT区域的数量;矩阵对角线处标注该基因组中发现的区域总数,以及未与Sscrofa11.1比对上的基因组区域数量。矩阵对角线上方区域展示基因组A中特有(未在基因组B中发现)的区域数量,反之亦然。参考基因组的缩写含义详见附加文件2表S5。 表S9。组装基因组中发现的所有非冗余NUMT区域列表,附带其在所有其他组装基因组中的存在/缺失情况。无法被定位至Sscrofa11.1基因组位置的NUMT区域未被纳入本列表。前六列包含唯一NUMT区域ID及其在Sscrofa11.1中的坐标;若NUMT区域在Sscrofa11.1中缺失,则标注其插入位点的坐标。剩余列对应组装基因组中唯一NUMT区域的相关信息:若有可用数据,每列将展示对应基因组中该NUMT区域的ID、坐标,以及其状态,状态定义如下: - CORRESPONDING(对应型):该基因组包含该NUMT区域; - COMPATIBLE SEQUENCE(序列兼容型):该基因组在NUMT区域坐标处与Sscrofa11.1存在比对,且另一基因组中的对应序列未被NUMT发现流程注释为NUMT; - PRIVATE(特有型):组装基因组与Sscrofa11.1的比对存在与两个基因组之一的NUMT区域坐标对应的间隙,意味着该NUMT插入仅发生在两个基因组之一中,因此该NUMT可被视为多态性位点。 每个状态列左侧为Sscrofa11.1的相关信息,右侧为其他基因组的对应信息。 表S10。通过挖掘全基因组测序(Whole Genome Sequencing, WGS)数据集鉴定到的所有新型NUMT列表。属于同一NUMT插入事件的断点将被赋予相同的NUMT区域ID。 表S11。基于WGS数据集估算的、各品种/种群及物种内新型NUMT区域携带者的频率。将报告每个NUMT区域的WGS数据集数量及其携带者占比。 表S12。通过聚合酶链式反应(Polymerase Chain Reaction, PCR)分析多个猪品种及种群中部分多态性NUMT得到的基因型信息。将报告携带三种基因型的猪只数量。 表S13。基于WGS数据集估算的、各品种/种群及物种内组装基因组中发现的所有非冗余NUMT区域携带者的频率。将报告每个NUMT区域的WGS数据集数量及其携带者占比。 表S14。来自组装基因组与WGS数据集的NUMT相关信息,包含年龄估算、在组装基因组与WGS数据集的存在/缺失情况,并按猪的物种及地理起源进行分类。 表S15。基于不同突变率,将NUMT区域划分为10–55百万年(Mya)及>55百万年两个年龄类别的比较年龄估算结果。 表S16。NUMT区域插入所在基因组区域的注释信息。




