遇见数据集

Mutation analysis of IGH associated with different isotypes

收藏
Figshare2014-04-07 更新2026-04-29 收录
官方服务:

资源简介:

Immunoglobulin heavy chain rearrangements were sequences from PBMCs of 8 individuals using IGHC primers for all isotypes bar IgE and IgD. Sequencing was carried out on the 454 pyrosequencing platform and rearranged immunoglobulin heavy chain sequences were partitioned against germline IGHV, IGHD and IGHJ using iHMMune-align. The datasets for each individual was filtered to remove non-productive rearrangements (frame shift of the IGHJ or stop codons), those with greater than 45 IGHV mutations, sequences which contained indels in the V or J, and those with ambiguities. Within each data only unique sequences were reatained (100% identical). Clonally-related sequences within each subject's dataset were also identified. Following germline reversion of all non-CDR3 nucleotide positions, sequence were clustered using the vmatch package (http://www.vmatch.de/) allowing up to 3 differences. A single representative for each cluster (representing a likely clonal set) was retained within the dataset. Representative sequences were selected based on the sequence with the highest copy number, or in cases where sequences shared the highest copy number, the least mutated sequence. If a cluster spanned multiple isotypes then a representative for each isotype was kept.

本数据集的免疫球蛋白重链重排序列取自8名个体的外周血单个核细胞(PBMCs),采用针对除免疫球蛋白ε(IgE)与免疫球蛋白δ(IgD)外所有同种型的免疫球蛋白重链(IGHC)引物进行扩增获取。测序在454焦磷酸测序平台上完成,随后使用iHMMune-align工具,将重排的免疫球蛋白重链序列与胚系IGHV、IGHD及IGHJ基因片段进行比对归类。对每个个体的数据集进行过滤处理:移除无功能重排序列(如IGHJ移码突变或存在终止密码子的序列)、IGHV突变数超过45的序列、V区或J区存在插入缺失(indel)的序列,以及存在碱基歧义的序列。每个数据集中仅保留完全一致的唯一序列(100%同源)。同时识别每个受试者数据集中的克隆相关序列。将所有非互补决定区3(CDR3)核苷酸位点回复至胚系序列后,使用vmatch软件包(http://www.vmatch.de/)对序列进行聚类,允许最多3个碱基差异。每个聚类(代表潜在的克隆组)仅保留一条代表性序列以纳入最终数据集。代表性序列的选取依据为拷贝数最高的序列;若多个序列拷贝数相同,则选取突变最少的序列。若一个聚类涵盖多种同种型,则为每种同种型分别保留一条代表性序列。

创建时间:
2014-04-07
二维码
社区交流群
二维码
科研交流群
商业服务