Pre-processed B cell receptor repertoire sequencing data from BioProject PRJNA527941
收藏资源简介:
<strong>Data Processing</strong> Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2). Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2) <strong>software_versions</strong> pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8 <strong>quality_thresholds</strong> FilterSeq.py pRESTO Q>20 <strong>paired_reads_assembly</strong> AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001 <strong>primer_match_cutoffs</strong> MaskPrimers.py pRESTO C primer & V primer maxerror 0.2 <strong>consensus_building</strong> BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5 <strong>collapsing_method</strong> CollapseSeq.py pRESTO <strong>germline_database </strong>IMGT <strong>Format</strong> Processed sequences are provided in a tab delimited file format, including the following annotations: <strong>C_CALL </strong>Isotype subclass <strong>SEQUENCE_ID </strong>Sequence identifier <strong>V_CALL </strong>V segment gene and allele <strong>D_CALL </strong>D segment gene and allele <strong>J_CALL </strong>J segment gene and allele <strong>JUNCTION_LENGTH </strong>Junction length <strong>CONSCOUNT </strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence. <strong>DUPCOUNT </strong>UMI count for the given unique sequence <strong>ISOTYPE </strong>Constant region primer (isotype) <strong>MU_COUNT_CDR_R </strong>Number of replacement mutations in CDR region <strong>MU_COUNT_CDR_S </strong>Number of silent mutations in CDR region <strong>MU_COUNT_FWR_R </strong>Number of replacement mutations in FWR region <strong>MU_COUNT_FWR_S </strong>Number of silent mutations in FWR region <strong>MUT_TOTAL </strong>Total number of mutations in V gene <strong>SEQUENCE_INPUT </strong>Full length sequence <strong>SEQUENCE_IMGT </strong>Gapped IMGT sequence <strong>V_GERM_START_VDJ </strong>position of the first nucleotide in ungapped V germline sequence alignment <strong>JUNCTION </strong>Junction nucleotide sequence <strong>GERMLINE_IMGT_D_MASK </strong>IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions <strong>Run </strong>ID of sequencing run <strong>Sample_type </strong>The tissue sampled (e.g Peripheral Blood, bone marrow, ..) <strong>Sex </strong>Sex of the Subject <strong>Age </strong>Age of the subject <strong>UNIQUE_ID </strong>Subject identifier <strong>SAMPLE_ID </strong>Sample identifier, linking back to raw data <strong>Subset </strong>Defined B cell subset <strong>Repertoire </strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG) <strong>R_SCDR </strong>R/S ratio in CDR region <strong>R_SFWR </strong>R/S ratio in FWR region <strong>V_FAM </strong>V family gene <strong>V_GENE </strong>V segment gene <strong>D_GENE </strong>D segment gene <strong>J_GENE </strong>J segment gene <strong>Clust_Rank </strong>Cluster rank <strong>Clust_REPRES </strong>Cluster representative <strong>Clust_SIZE </strong>Cluster size <strong>Clust_MAXFREQ </strong>Cluster maximum frequency <strong>Clust_SHAREDNESS </strong>Cluster sharedness <strong>CDR3_AA_GRAVY </strong>CDR3 hydrophobicity index <strong>CDR3_AA_CHARGE </strong>CDR3 charge <strong>CDRH3PDB </strong>CDRH3 PDB (Structure) code <strong>H1Canon </strong>H1 Canonical class <strong>H2Canon </strong>H2 Canonical class <strong>H1_GERMLINE </strong>H1 Germline Canonical class <strong>H2_GERMLINE </strong>H2 Germline Canonical class <strong>References</strong> 1. Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O’Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein. 2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires. <em>Bioinformatics</em>30: 1930–1932. 2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein. 2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data. <em>Bioinformatics</em>31: 3356–3358. 3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool. <em>Nucleic Acids Res.</em>41. 4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads. <em>Genome Res.</em>21: 936–939.
数据处理 样本通过Illumina索引进行解复用,并使用Immcantation工具包(Immcantation toolkit)进行处理(1,2)。原始FASTQ文件基于质量阈值20进行过滤。当配对读段的最小长度为10 nt、最大错误率为0.3且显著性阈值为0.0001时,将其进行拼接。具有相同唯一分子标签(Unique Molecular Identifier, UMI)的读段将被合并为一致序列。全长序列一致、恒定引物相同但UMI不同的读段将被进一步合并。随后将序列提交至IgBlast(IgBlast)进行VDJ基因分型和序列注释。使用Stampy(Stampy)将恒定区序列比对至种系序列。使用shazam R包计算V基因的突变数量与类型(2)。 软件版本 pRESTO:0.5.3、Change-O:0.3.4、IgBlast 1.6.1、Stampy 1.0.21、shazam 0.1.8 质量阈值 FilterSeq.py(pRESTO):Q>20 配对读段组装 AssemblePairs.py(pRESTO):最小长度(minlen)10 nt、最大错误率(maxerror)0.3、显著性阈值(alpha)0.0001 引物匹配截断值 MaskPrimers.py(pRESTO):C引物与V引物的最大错误率为0.2 一致序列构建 BuildConsensus.py(pRESTO):最大错误率0.1、最大间隙率0.5 合并方法 CollapseSeq.py(pRESTO) 种系数据库 IMGT种系数据库 数据格式 处理后的序列以制表符分隔文件格式提供,包含以下注释字段: C_CALL:同种型亚类 SEQUENCE_ID:序列标识符 V_CALL:V区段基因与等位基因 D_CALL:D区段基因与等位基因 J_CALL:J区段基因与等位基因 JUNCTION_LENGTH:连接区长度 CONSCOUNT:生成UMI一致序列的原始读段计数,针对给定唯一序列的所有UMI求和 DUPCOUNT:给定唯一序列的UMI计数 ISOTYPE:恒定区引物(同种型) MU_COUNT_CDR_R:CDR区的替换突变数 MU_COUNT_CDR_S:CDR区的沉默突变数 MU_COUNT_FWR_R:FWR区的替换突变数 MU_COUNT_FWR_S:FWR区的沉默突变数 MUT_TOTAL:V基因总突变数 SEQUENCE_INPUT:全长序列 SEQUENCE_IMGT:带缺口的IMGT序列 V_GERM_START_VDJ:无缺口V种系序列比对的第一个核苷酸位置 JUNCTION:连接区核苷酸序列 GERMLINE_IMGT_D_MASK:IMGT带缺口的种系核苷酸序列,其中NP1-D-NP2区域以Ns标记 Run:测序运行ID Sample_type:采样组织(如外周血、骨髓等) Sex:受试者性别 Age:受试者年龄 UNIQUE_ID:受试者标识符 SAMPLE_ID:样本标识符,可关联至原始数据 Subset:定义的B细胞亚群 Repertoire:定义的B细胞免疫组库(初始型、记忆型IgM/IgD、IgA、IgG) R_SCDR:CDR区的R/S比值 R_SFWR:FWR区的R/S比值 V_FAM:V家族基因 V_GENE:V区段基因 D_GENE:D区段基因 J_GENE:J区段基因 Clust_Rank:聚类排名 Clust_REPRES:聚类代表序列 Clust_SIZE:聚类大小 Clust_MAXFREQ:聚类最大频率 Clust_SHAREDNESS:聚类共享性 CDR3_AA_GRAVY:CDR3氨基酸疏水性指数 CDR3_AA_CHARGE:CDR3氨基酸电荷值 CDRH3PDB:CDRH3的PDB(结构)代码 H1Canon:H1经典型类别 H2Canon:H2经典型类别 H1_GERMLINE:H1种系经典型类别 H2_GERMLINE:H2种系经典型类别 参考文献 1. Vander Heiden J A, Yaari G, Uduman M, Stern J N H, O'Connor K C, Hafler D A, Vigneault F, Kleinstein S H. 2014. PRESTO:用于处理淋巴细胞受体免疫组库高通量测序原始读段的工具包. *Bioinformatics*, 30: 1930–1932. 2. Gupta N T, Vander Heiden J A, Uduman M, Gadala-Maria D, Yaari G, Kleinstein S H. 2015. Change-O:用于分析大规模B细胞免疫球蛋白免疫组库测序数据的工具包. *Bioinformatics*, 31: 3356–3358. 3. Ye J, Ma N, Madden T L, Ostell J M. 2013. IgBLAST:免疫球蛋白可变区序列分析工具. *Nucleic Acids Research*, 41. 4. Lunter G, Goodson M. 2011. Stampy:用于灵敏快速比对Illumina序列读段的统计算法. *Genome Research*, 21: 936–1939.



