遇见数据集

Dataset for: Pre-pandemic artificial MERS analog of polyfunctional SARS-CoV-2 S1/S2 furin cleavage site domain is unique among spike proteins of genus Betacoronavirus

收藏
Zenodo2024-12-23 更新2026-05-26 收录
官方服务:

资源简介:

Data File Descriptions and Methods Data file 1 [betacov_matching_IPR042578.fasta]: Representative set of 2,465 betacoronavirus S protein overlapping homologous superfamily sequences retrieved in fasta format on 4 December 2022 from the InterPro repository at https://www.ebi.ac.uk/interpro/entry/InterPro/IPR042578/. Data File 2 [betacov_matching_IPR042578_motif.fasta]: With Data File 1 as input, extracted 98,122 furin cleavage site (FCS) output motifs of 20 amino acids length, including overlapping and redundant sequences, produced with the FindFur algorithm with preset parameters as described by (Gu, 2020). FindFur as used was deposited on 15 December 2020 at the GitHub software repository at https://github.com/chwisteeng/FindFur. Data File 3 [table_s1s2_hits_betacov_polyf.pdf]: Compiled summary table of sequence hits (PDF) of spike S1/S2 domains across genus Betacoronavirus. The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences. Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS. Data File 4 [table_s1s2_hits_betacov_polyf.xlsx]: Compiled summary table of sequence hits (MS Excel) of spike S1/S2 domains across genus Betacoronavirus. The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences. Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS. Data File 5 [betacov_s1s2_nls_pat7_furin_psort.txt]: Nuclear localization signal (NLS) detection output for 5 representative betacoronavirus spike sequence domains, including the positive hits for pat7 in SARS-CoV-2 and for MERS-MA30 CoV. NLS predictions used the PSORT algorithm available as a webservice at https://wolfpsort.hgc.jp/ which is based on the work of Nakai and Horton (Nakai and Horton, 1999). Numbering refers to Data File 3 and Data File 4. Data File 6 [betacov_s1s2_oglyc_netogly.txt]: Detection output for 5 representative betacoronavirus spike sequence domains tested for Thr/Ser O-glycosite residue pairs with the standard prediction software NetOGlyc4.0 (Steentoft et al., 2013) as available at https://services.healthtech.dtu.dk/services/NetOGlyc-4.0/. Positive hits have scores above 0.5. Numbering refers to Data File 3 and Data File 4. Data File 7 [betacov_s1s2_nls_pat7_furin_blastp.txt]: Comprehensive sequence database searches using were performed using the NCBI protein BLAST (blastp) algorithm with webservice available at https://blast.ncbi.nlm.nih.gov/Blast.cgi?PAGE=Proteins. The following blastp search parameters and settings were used: Word size=2; Expect value=200000; Hitlist size=500; Gapcosts=9,1; Matrix=PAM30; Filter string=F; Genetic Code=1;Window Size=40; Threshold=11; Composition-based stats=0; Database Posted date=Jan 19, 2023 2:59 AM; Number of letters=17,117,563; Number of sequences=10,766; Entrez query: Includes: Betacoronavirus (taxid:694002); Excludes: SARS-CoV-2 (taxid:2697049). The six polyfunctional input query consensus motif sequences were TXXPR(K/H/R)XRSX and TXXPRX(K/H/R)RSX. References Gu, C., 2020. FindFur: A Tool for Predicting Furin Cleavage Sites of Viral Envelope Substrates. Master’s Thesis, San Jose State University, CA, USA. doi: 10.31979/etd.4ahv-9jya Gangavarapu K, Latif AA, Mullen JL, Alkuzweny M, Hufbauer E, Tsueng G, Haag E, Zeller M, Aceves CM, Zaiets K, Cano M, Zhou X, Qian Z, Sattler R, Matteson NL, Levy JI, Lee RTC, Freitas L, Maurer-Stroh S; GISAID Core and Curation Team; Suchard MA, Wu C, Su AI, Andersen KG, Hughes LD. Outbreak.info genomic reports: scalable and dynamic surveillance of SARS-CoV-2 variants and mutations. Nat Methods. 2023. 20(4):512-522. doi: 10.1038/s41592-023-01769-3. Nakai, K., Horton, P., 1999. PSORT: a program for detecting sorting signals in proteins and predicting their subcellular localization. Trends Biochem Sci 24, 34–36. doi: 10.1016/s0968-0004(98)01336-x Steentoft, C., Vakhrushev, S.Y., Joshi, H.J., Kong, Y., Vester-Christensen, M.B., Schjoldager, K.T.-B.G., Lavrsen, K., Dabelsteen, S., Pedersen, N.B., Marcos-Silva, L., Gupta, R., Bennett, E.P., Mandel, U., Brunak, S., Wandall, H.H., Levery, S.B., Clausen, H., 2013. Precision mapping of the human O-GalNAc glycoproteome through SimpleCell technology. EMBO J 32, 1478–1488. doi: 10.1038/emboj.2013.79

数据文件描述与实验方法 数据文件1 [betacov_matching_IPR042578.fasta]:2022年12月4日从欧洲生物信息研究所的InterPro数据库(https://www.ebi.ac.uk/interpro/entry/InterPro/IPR042578/)获取的2465条β冠状病毒刺突(S)蛋白重叠同源超家族序列的代表性集合,格式为FASTA。 数据文件2 [betacov_matching_IPR042578_motif.fasta]:以数据文件1作为输入,通过FindFur工具(预设参数遵循Gu, 2020所述方法)提取得到98122条长度为20个氨基酸的弗林蛋白酶切割位点(Furin Cleavage Site, FCS)基序序列,包含重叠及冗余序列。FindFur工具于2020年12月15日上传至GitHub软件仓库,地址为https://github.com/chwisteeng/FindFur。 数据文件3 [table_s1s2_hits_betacov_polyf.pdf]:β冠状病毒属刺突S1/S2结构域序列命中结果的汇总统计表(PDF格式)。该汇总表对数据文件2中的序列进行了筛选:移除对应于刺突蛋白片段(即GenBank提交的不完整长度刺突蛋白)的序列以及重复序列(即在20氨基酸基序窗口内完全重叠的冗余片段),并为多组完全一致的序列选取一条代表性序列。序列的采集日期与地理信息从NCBI GenBank蛋白质数据库(https://www.ncbi.nlm.nih.gov/protein/)获取。对于SARS-CoV-2刺突蛋白变异株,这些数据还通过SARS-CoV-2谱系突变追踪工具(Gangavarapu, 2023)进行了交叉验证,该工具可通过https://outbreak.info访问,其数据基于全球流感数据共享全球倡议(Global Initiative on Sharing All Influenza Data, GISAID)的大规模测序数据,GISAID官网地址为https://gisaid.org/。图表中,实线(-)表示pat7核定位信号(Nuclear Localization Signal, NLS),星号(*)表示O-糖基化位点,脱字符(^)表示弗林蛋白酶切割位点(FCS)。 数据文件4 [table_s1s2_hits_betacov_polyf.xlsx]:β冠状病毒属刺突S1/S2结构域序列命中结果的汇总统计表(Microsoft Excel格式)。该汇总表对数据文件2中的序列进行了筛选:移除对应于刺突蛋白片段(即GenBank提交的不完整长度刺突蛋白)的序列以及重复序列(即在20氨基酸基序窗口内完全重叠的冗余片段),并为多组完全一致的序列选取一条代表性序列。序列的采集日期与地理信息从NCBI GenBank蛋白质数据库(https://www.ncbi.nlm.nih.gov/protein/)获取。对于SARS-CoV-2刺突蛋白变异株,这些数据还通过SARS-CoV-2谱系突变追踪工具(Gangavarapu, 2023)进行了交叉验证,该工具可通过https://outbreak.info访问,其数据基于全球流感数据共享全球倡议(GISAID)的大规模测序数据,GISAID官网地址为https://gisaid.org/。图表中,实线(-)表示pat7核定位信号(NLS),星号(*)表示O-糖基化位点,脱字符(^)表示弗林蛋白酶切割位点(FCS)。 数据文件5 [betacov_s1s2_nls_pat7_furin_psort.txt]:5条代表性β冠状病毒刺突序列结构域的核定位信号(NLS)检测结果,其中包含SARS-CoV-2及MERS-MA30 CoV的pat7阳性命中结果。核定位信号预测采用PSORT算法,该算法作为Web服务可通过https://wolfpsort.hgc.jp/访问,其开发基于Nakai与Horton(1999)的研究工作。序列编号对应数据文件3与数据文件4。 数据文件6 [betacov_s1s2_oglyc_netogly.txt]:针对5条代表性β冠状病毒刺突序列结构域的苏氨酸/丝氨酸O-糖基化位点残基对的检测结果,检测采用标准预测软件NetOGlyc4.0(Steentoft et al., 2013),该软件可通过https://services.healthtech.dtu.dk/services/NetOGlyc-4.0/访问。阳性命中结果的评分高于0.5。序列编号对应数据文件3与数据文件4。 数据文件7 [betacov_s1s2_nls_pat7_furin_blastp.txt]:采用NCBI蛋白质BLAST(blastp)算法进行的全面序列数据库搜索,该算法的Web服务可通过https://blast.ncbi.nlm.nih.gov/Blast.cgi?PAGE=Proteins访问。本次搜索采用的参数与设置如下:字长(Word size)=2;期望值(Expect value)=200000;命中列表大小(Hitlist size)=500;空位代价(Gapcosts)=9,1;矩阵(Matrix)=PAM30;过滤字符串(Filter string)=F;遗传密码(Genetic Code)=1;窗口大小(Window Size)=40;阈值(Threshold)=11;基于组成的统计(Composition-based stats)=0;数据库发布日期=2023年1月19日 2:59;字母序列总数=17,117,563;序列总数=10,766;Entrez查询条件:包含:β冠状病毒属(分类学ID:694002);排除:SARS-CoV-2(分类学ID:2697049)。本次搜索使用的6条多功能输入查询共有基序序列为TXXPR(K/H/R)XRSX与TXXPRX(K/H/R)RSX。 参考文献 [1] Gu, C., 2020. FindFur:病毒包膜底物弗林蛋白酶切割位点预测工具. 硕士学位论文, 美国加州圣何塞州立大学. doi: 10.31979/etd.4ahv-9jya [2] Gangavarapu K, Latif AA, Mullen JL, Alkuzweny M, Hufbauer E, Tsueng G, Haag E, Zeller M, Aceves CM, Zaiets K, Cano M, Zhou X, Qian Z, Sattler R, Matteson NL, Levy JI, Lee RTC, Freitas L, Maurer-Stroh S; GISAID核心与审核团队; Suchard MA, Wu C, Su AI, Andersen KG, Hughes LD. Outbreak.info基因组报告:SARS-CoV-2变异株与突变的可扩展动态监测. 《自然-方法》, 2023, 20(4):512-522. doi: 10.1038/s41592-023-01769-3 [3] Nakai, K., Horton, P., 1999. PSORT:一款用于检测蛋白质分选信号并预测其亚细胞定位的程序. 《生物化学趋势》, 24, 34–36. doi: 10.1016/s0968-0004(98)01336-x [4] Steentoft, C., Vakhrushev, S.Y., Joshi, H.J., Kong, Y., Vester-Christensen, M.B., Schjoldager, K.T.-B.G., Lavrsen, K., Dabelsteen, S., Pedersen, N.B., Marcos-Silva, L., Gupta, R., Bennett, E.P., Mandel, U., Brunak, S., Wandall, H.H., Levery, S.B., Clausen, H., 2013. 基于SimpleCell技术精准绘制人类O-GalNAc糖蛋白组. 《欧洲分子生物学组织杂志》, 32, 1478–1488. doi: 10.1038/emboj.2013.79

提供机构:
Zenodo
创建时间:
2024-08-01
二维码
社区交流群
二维码
科研交流群
商业服务