遇见数据集

Microbial Recombination with Population Structure

收藏
NIAID Data Ecosystem2026-03-12 收录
数据链接:
官方服务:

资源简介:

This collection of data files contains, for each bacterial species: All raw genome sequence files. The core genome alignment obtained with REALPHY (this is the file with the .phy extension). A file with all SNP columns in the core genome alignment (the file name starts with columns_). A file listing all SNP types sorted from most to least common (the file name starts with snp_stats_) Two files containing the results of the pairwise analysis. First, a file with, for each pair, the histogram of SNP counts per alignment block (the file name ends in _histograms). And second, a file with the results of the mixture modeling (the file with the .pkl extension). In addition, for M. tuberculosis there is a subfolder with information about which strains have since been retracted from the database. The formats of these files are as follows: The raw genome sequence files are in FASTA format (.fasta or .fna). The core genome alignment is in PHYLIP multiple alignment format (.phy). The snp_stats file starts with a header line listing the total number of columns in the alignment with 1, 2, 3, and 4 different nucleotides. Each next line in the file corresponds to an observed SNP-type, sorted from most to least common. Each SNP line has the following columns: The total amount of genomic DNA associated with these SNP columns (associating each conserved alignment column to its closest SNP). The total number of occurrences of this SNP type. The number of strains sharing the minority allele. A bit-pattern describing the SNP type, with 1 for the strains sharing the minority allele, and zero for the others. The strains are sorted in the same order as in the PHYLIP alignment file. A list of all the strains sharing the minority allele. The columns_ file has one SNP per line, giving the position in the alignment plus the bit-pattern describing the SNP. The _histogram file contains, for each pair of strains, a histogram counting the number of 1Kb blocks with 0, 1, 2, etc SNPs. Note that these counts come from 1 kilobase sliding windows along the core genome alignment, sliding the window by 100 bases at a time, i.e. an alignment column will typically occur in 10 blocks. A pickle file with, for each pair, the results of the mixture modeling. Each line corresponds to a pair and has these fields: [spec1, spec2, div, Lpois, r_nomix, Lmix, rho, r, a, lam, mutpois, mutrecomb, cut] which correspond to: Name of strain 1. Name of strain 2. Their overall nucleotide divergence. The log-likelihood under a model assuming SNP counts form a simple Poisson distribution. The parameter of this fitted Poisson distribution. The log-likelihood of the mixture of a Poisson and negative binomial The fraction rho assigned to the Poisson part of the mixture. The parameter of the Poisson component. The exponent a of the negative binomial component. The second parameter (lambda) of the negative binomial. The estimated total number of mutations in the Poisson component. The estimated total number of mutations in the negative binomial component. The value at which the likelihood of negative poisson component starts exceeding the likelihood of the negative binomial component. In addition, for the human data we provide a PHYLIP multiple genome alignment and a file with all SNP columns.

本数据集的数据文件集合涵盖了针对每一种细菌物种的如下内容: 1. 全部原始基因组序列文件。 2. 通过REALPHY工具得到的核心基因组比对文件(扩展名为.phy)。 3. 包含核心基因组比对中所有单核苷酸多态性(Single Nucleotide Polymorphism, SNP)列的文件(文件名以columns_开头)。 4. 按出现频率从高到低排序的所有SNP类型统计文件(文件名以snp_stats_开头)。 5. 两份成对分析结果文件:其一为记录每对菌株的比对区块SNP计数直方图的文件(文件名以_histograms结尾);其二为混合模型分析结果文件(扩展名为.pkl)。 此外,针对结核分枝杆菌(M. tuberculosis),还设有一个子文件夹,内含有关哪些菌株已从数据库中撤回的相关信息。 各文件的格式规范如下: 1. 原始基因组序列文件采用FASTA格式(扩展名:.fasta 或 .fna)。 2. 核心基因组比对文件采用PHYLIP多序列比对格式(扩展名:.phy)。 3. snp_stats_文件的首行为表头,统计了具有1、2、3、4种不同核苷酸的比对列总数。文件后续每一行对应一种观测到的SNP类型,按出现频率从高到低排序。每行SNP数据包含以下列: - 与这些SNP列相关的总基因组DNA量(将每个保守比对列与其最邻近的SNP进行关联)。 - 该SNP类型的总出现次数。 - 携带次要等位基因的菌株数量。 - 描述该SNP类型的位模式:携带次要等位基因的菌株标记为1,其余菌株标记为0。菌株的排序与PHYLIP比对文件中的顺序保持一致。 - 所有携带次要等位基因的菌株列表。 4. columns_文件每行对应一个SNP,给出其在比对中的位置以及描述该SNP的位模式。 5. _histograms文件包含每一对菌株的直方图数据,用于统计每1Kb滑动窗口内含有0、1、2等数量SNP的区块数。请注意,这些统计值源自沿核心基因组比对的1kb滑动窗口,每次滑动100个碱基,即单个比对列通常会出现在10个区块中。 6. pickle格式文件存储了每对菌株的混合模型分析结果。每行对应一组菌株对,包含以下字段:[spec1, spec2, div, Lpois, r_nomix, Lmix, rho, r, a, lam, mutpois, mutrecomb, cut],各字段对应含义如下: - 菌株1的名称。 - 菌株2的名称。 - 二者的整体核苷酸差异度。 - 假设SNP计数服从简单泊松分布的模型下的对数似然值。 - 该拟合泊松分布的参数。 - 泊松分布与负二项分布混合模型的对数似然值。 - 混合模型中泊松部分所占比例ρ。 - 泊松分量的分布参数。 - 负二项分量的形状参数a。 - 负二项分布的第二个参数λ。 - 泊松分量下估计的总突变数。 - 负二项分量下估计的总突变数。 - 负泊松分量似然超过负二项分量似然的临界值。 此外,针对人类相关数据集,我们还提供了一份PHYLIP格式的多基因组比对文件,以及一份包含所有SNP列的文件。

创建时间:
2021-01-07
二维码
社区交流群
二维码
科研交流群
商业服务