遇见数据集

Pre-processed B cell receptor repertoire sequencing data from BioProject PRJNA527941

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Data Processing</strong> Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2). Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2) <strong>software_versions</strong> pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8 <strong>quality_thresholds</strong> FilterSeq.py pRESTO Q&gt;20 <strong>paired_reads_assembly</strong> AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001 <strong>primer_match_cutoffs</strong> MaskPrimers.py pRESTO C primer &amp; V primer maxerror 0.2 <strong>consensus_building</strong> BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5 <strong>collapsing_method</strong> CollapseSeq.py pRESTO <strong>germline_database </strong>IMGT <strong>Format</strong> Processed sequences are provided in a tab delimited file format, including the following annotations: <strong>C_CALL </strong>Isotype subclass <strong>SEQUENCE_ID </strong>Sequence identifier <strong>V_CALL </strong>V segment gene and allele <strong>D_CALL </strong>D segment gene and allele <strong>J_CALL </strong>J segment gene and allele <strong>JUNCTION_LENGTH </strong>Junction length <strong>CONSCOUNT </strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence. <strong>DUPCOUNT </strong>UMI count for the given unique sequence <strong>ISOTYPE </strong>Constant region primer (isotype) <strong>MU_COUNT_CDR_R </strong>Number of replacement mutations in CDR region <strong>MU_COUNT_CDR_S </strong>Number of silent mutations in CDR region <strong>MU_COUNT_FWR_R </strong>Number of replacement mutations in FWR region <strong>MU_COUNT_FWR_S </strong>Number of silent mutations in FWR region <strong>MUT_TOTAL </strong>Total number of mutations in V gene <strong>SEQUENCE_INPUT </strong>Full length sequence <strong>SEQUENCE_IMGT </strong>Gapped IMGT sequence <strong>V_GERM_START_VDJ </strong>position of the first nucleotide in ungapped V germline sequence alignment <strong>JUNCTION </strong>Junction nucleotide sequence <strong>GERMLINE_IMGT_D_MASK </strong>IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions <strong>Run </strong>ID of sequencing run <strong>Sample_type </strong>The tissue sampled (e.g Peripheral Blood, bone marrow, ..) <strong>Sex </strong>Sex of the Subject <strong>Age </strong>Age of the subject <strong>UNIQUE_ID </strong>Subject identifier <strong>SAMPLE_ID </strong>Sample identifier, linking back to raw data <strong>Subset </strong>Defined B cell subset <strong>Repertoire </strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG) <strong>R_SCDR </strong>R/S ratio in CDR region <strong>R_SFWR </strong>R/S ratio in FWR region <strong>V_FAM </strong>V family gene <strong>V_GENE </strong>V segment gene <strong>D_GENE </strong>D segment gene <strong>J_GENE </strong>J segment gene <strong>Clust_Rank </strong>Cluster rank <strong>Clust_REPRES </strong>Cluster representative <strong>Clust_SIZE </strong>Cluster size <strong>Clust_MAXFREQ </strong>Cluster maximum frequency <strong>Clust_SHAREDNESS </strong>Cluster sharedness <strong>CDR3_AA_GRAVY </strong>CDR3 hydrophobicity index <strong>CDR3_AA_CHARGE </strong>CDR3 charge <strong>CDRH3PDB </strong>CDRH3 PDB (Structure) code <strong>H1Canon </strong>H1 Canonical class <strong>H2Canon </strong>H2 Canonical class <strong>H1_GERMLINE </strong>H1 Germline Canonical class <strong>H2_GERMLINE </strong>H2 Germline Canonical class <strong>References</strong> 1. Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O’Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein. 2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires. <em>Bioinformatics</em>30: 1930–1932. 2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein. 2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data. <em>Bioinformatics</em>31: 3356–3358. 3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool. <em>Nucleic Acids Res.</em>41. 4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads. <em>Genome Res.</em>21: 936–939.

**数据处理** 样本通过Illumina索引进行解复用,并使用Immcantation工具包[1,2]进行处理。原始fastq文件基于质量得分阈值20进行过滤。双端测序reads在满足最小长度10 nt、最大错误率0.3以及显著性阈值0.0001的条件下进行拼接。具有相同唯一分子标识符(Unique Molecular Identifier, UMI)的reads将被合并为一致序列;具有完全一致的全长序列、相同的恒定区引物但UMI不同的reads将进一步合并。随后将序列提交至IgBLAST[3]进行VDJ基因分配与序列注释。恒定区序列使用Stampy工具[4]比对至种系参考序列。使用shazam R包[2]计算V基因的突变数量与类型。 **软件版本** pRESTO:0.5.3,Change-O:0.3.4,IgBLAST 1.6.1,Stampy 1.0.21,shazam 0.1.8 **质量过滤阈值** FilterSeq.py pRESTO Q>20 **双端reads组装参数** AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001 **引物匹配阈值** MaskPrimers.py pRESTO C引物与V引物maxerror 0.2 **一致序列构建参数** BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5 **合并方法** CollapseSeq.py pRESTO **种系参考数据库** IMGT **数据格式** 经处理的序列以制表符分隔的文件格式提供,包含以下注释信息: **C_CALL**:同种型亚型 **SEQUENCE_ID**:序列标识符 **V_CALL**:V区段基因及等位基因 **D_CALL**:D区段基因及等位基因 **J_CALL**:J区段基因及等位基因 **JUNCTION_LENGTH**:连接区长度 **CONSCOUNT**:生成UMI一致序列的原始reads计数,针对给定唯一序列的所有UMI进行求和 **DUPCOUNT**:给定唯一序列的UMI计数 **ISOTYPE**:恒定区引物(同种型) **MU_COUNT_CDR_R**:CDR区的替换突变数 **MU_COUNT_CDR_S**:CDR区的沉默突变数 **MU_COUNT_FWR_R**:FWR区的替换突变数 **MU_COUNT_FWR_S**:FWR区的沉默突变数 **MUT_TOTAL**:V基因总突变数 **SEQUENCE_INPUT**:全长序列 **SEQUENCE_IMGT**:带缺口的IMGT序列 **V_GERM_START_VDJ**:无缺口的V种系序列比对中第一个核苷酸的位置 **JUNCTION**:连接区核苷酸序列 **GERMLINE_IMGT_D_MASK**:IMGT带缺口的种系核苷酸序列,其中NP1-D-NP2区域以Ns标记 **Run**:测序运行ID **Sample_type**:采样组织类型(例如外周血、骨髓等) **Sex**:受试者性别 **Age**:受试者年龄 **UNIQUE_ID**:受试者标识符 **SAMPLE_ID**:样本标识符,可关联至原始测序数据 **Subset**:定义的B细胞亚群 **Repertoire**:定义的B细胞免疫组库(初始型、记忆型IgM/IgD、IgA、IgG) **R_SCDR**:CDR区的R/S比值 **R_SFWR**:FWR区的R/S比值 **V_FAM**:V家族基因 **V_GENE**:V区段基因 **D_GENE**:D区段基因 **J_GENE**:J区段基因 **Clust_Rank**:聚类排名 **Clust_REPRES**:聚类代表序列 **Clust_SIZE**:聚类大小 **Clust_MAXFREQ**:聚类最大频率 **Clust_SHAREDNESS**:聚类共享性 **CDR3_AA_GRAVY**:CDR3氨基酸疏水性指数 **CDR3_AA_CHARGE**:CDR3氨基酸电荷值 **CDRH3PDB**:CDRH3的PDB(结构)编号 **H1Canon**:H1经典型类别 **H2Canon**:H2经典型类别 **H1_GERMLINE**:H1种系经典型类别 **H2_GERMLINE**:H2种系经典型类别 **参考文献** 1. Vander Heiden J A, Yaari G, Uduman M, Stern J N H, O'Connor K C, Hafler D A, Vigneault F, Kleinstein S H. 2014. PRESTO:用于处理淋巴细胞受体组库高通量测序原始reads的工具包. *Bioinformatics*, 30: 1930–1932. 2. Gupta N T, Vander Heiden J A, Uduman M, Gadala-Maria D, Yaari G, Kleinstein S H. 2015. Change-O:用于分析大规模B细胞免疫组库测序数据的工具包. *Bioinformatics*, 31: 3356–3358. 3. Ye J, Ma N, Madden T L, Ostell J M. 2013. IgBLAST:免疫球蛋白可变区序列分析工具. *Nucleic Acids Res.*, 41. 4. Lunter G, Goodson M. 2011. Stampy:用于灵敏快速比对Illumina测序reads的统计算法. *Genome Res.*, 21: 936–939.

提供机构:
Zenodo
创建时间:
2019-04-15
二维码
社区交流群
二维码
科研交流群
商业服务