遇见数据集

DNA loss model explains the evolution of the neuropeptide LWamide, APGWamide, APGW/AKH, RPCH, AKH, ACP, CRZ, and GnRH families

收藏
Zenodo2023-06-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>R1: Establishment and purification of neuropeptide sequences</strong> The LW, APGW, RPCH, AKH, CRZ, and GnRH neuropeptide families were searched in the GenBank database using 10 keywords: the neuropeptide name, the precursor abbreviation, the full name of the precursor, the full name of the precursor with the word “prepropeptide,” and the combinations of these terms. The candidate sequences were downloaded in FASTA format using the appropriate commands in the GenBank database. The AKH neuropeptide family was classified according to the groups published in the literature, as well as the amino acid number and sequence. Furthermore, the ACP hybrid family was identified in the GenBank database using BLAST alignments. <strong>C00: Neuropeptide Precursor. </strong>Eight folders were named with the initials of each neuropeptide family. The AKH family folder was the only one containing four subfolders. All of the folders contained the same type of files: three text files named after the neuropeptide initials and the obtained result. The files identified with the words “<em>with codes</em>” contained the sequences with the codes generated for this study, whereas the documents with the word “<em>Full</em>” contained the GenBank database search results obtained with the 10 aforementioned keywords. These files were located in a folder named “<em>Fasta Keywords.</em>” Each file contained the results from each respective keyword. The files with the words “<em>selected EA</em>” contained the sequences that were selected for evolutionary analyses. <strong>C01: BLAST ACP</strong>. The text file named “00 BLAST ACP” contains the BLAST alignment results obtained from the NCBI database generated with the Adipokinetic Hormone/Corazonin-related peptide from the transcriptome of <em>Callinectes toxotes</em>. The file named “01 ACP Selected” contains the precursors selected for this study. All sequences were in FASTA format and contained the codes summarized in Supplementary Material 3 “<em>Database Sequences.</em>” The file named “<em>02 ACP selected EA</em>” contains the ACP precursors of other species, which were used for the evolutionary analyses of <em>C. toxotes</em> ACP. The PDF file titled “<em>03 ACP ProP 1.0 Serv</em>” contains the results of the proteolytic cleavage sites of the precursors indicated in the file named “<em>02 ACP selected EA,</em>” which were generated using the aforementioned software. <strong>C02: BLAST VP.</strong> The folder contains the results of the BLAST alignment against the NCBI database, which were generated with the virtual peptide sequences reported by Martinez-Perez et al. (2007). This folder contains seven text files. The name of each file corresponds to the precursor and species in which it was identified. Moreover, the PDF document named “<em>Virtual peptides ProP 1.0 Serv</em>” contains the results of the proteolytic cleavage sites generated with the aforementioned software. <strong>C03: Debugging sequences with software.</strong> This folder contains three subfolders containing the results obtained with each software used in this study for the detection of each of the neuropeptide sequences using the appropriate keywords. The folder named “<em>BioDataToolKit</em>” contains six subfolders with the abbreviated name of each neuropeptide. Additionally, there is a file containing the sequences downloaded from the GenBank database, as well as a Microsoft Excel file containing the details generated by the software. The name of each file corresponds to the keywords used for each search. The software used in this study can be found in the following repository: https://github.com/rduarte24/BiodataToolkit. The folder named “<em>Pro1.0Server</em>” was organized in the same way as the results derived for the “<em>BioDataToolKit</em>” for each neuropeptide family. However, each of the neuropeptide folders contained a file with the pertinent sequences whereas another file contained the endoproteolytic cleavage sites of the neuropeptide precursors obtained with the software. The folder named “Proteios” contains seven files. The file names indicate the precursor analyzed with the software and the identified sequences in FASTA format. The Proteios software is available in the following website: https://github.com/Martin-Munive/Proteios. <strong>C04: Neuropeptide precursors for evolutionary analysis.</strong> Files with the sequences of the neuropeptide precursors used for the generation of the phylogenetic trees in Supplementary Materials 4 and 7. The name of each file corresponds to the name of each of the analyzed neuropeptides. <strong>R2: Transcriptome BLAST</strong> Microsoft Excel file containing the BLAST alignments conducted using the sequences of the AKH/CRZ-related peptide (ACP) from <em>C. toxotes</em> and Corazonin (CRZ) from <em>C. arcuatus</em>. The following information is summarized in the spreadsheets named <em>C. toxotes</em> and <em>C. arcuatus</em>: Column A, neuropeptide name; Column B, species name; Columns C–G, BLAST alignment results; Column H, GenBank protein accession number; Column I, precursor sequence. <strong>R3: </strong><strong>Construction of neuropeptide database</strong> Microsoft Excel file with information pertaining to the database and a detailed description of each of the neuropeptide precursors analyzed in this study. The Excel file contains seven spreadsheet tabs. Each of the tabs contains the following columns: <strong>Neuropeptides.</strong> Column A, sequence numbering in descending order; Column B, neuropeptide name; Column C, identification code used in this study; Column D, accession number; Columns E–G, species taxonomy; Columns H–L, GenBank sequence description; Columns M–N, literature reference and link. <strong>Taxonomy.</strong> Taxonomic description of each of the examined species derived from the NCBI database. <strong>Sequences evolutionary anal</strong>. This tab contains the code developed for this work in Column C; the GenBank accession codes of each neuropeptide are summarized in Column D and species taxonomy details are summarized in Columns E y F. <strong>Table of differences.</strong> Column B shows the codes of identical sequences and Column C shows the code of the sequence selected for this study. <strong>Codes deleted. </strong>This tab contains the accession codes of the species and the species name but contains no details on the properties of the neuropeptide precursors. <strong>Sequences Paper</strong>. Neuropeptide sequences reported in previous studies that were later reported in the GenBank database. The sequences marked with asterisks have not been previously reported in public databases. The codes used in this study to designate the sequences are also included. <strong>Keywords. </strong>Keywords used to conduct the GenBank database searches to obtain the members of each neuropeptide family. <strong>R4: <em>In silico</em> validation, alignments, and phylogenetic relationships</strong> Generated phylogenetic trees and results obtained from individual runs for each of the neuropeptide families with the DNA-LM and Kalign parameters using the IQ-TREE software. The folder named “<em>RUN</em>” contains the “<em>DNALM and kalign 2.0 default parameters</em>” subfolder. Both folders contain 11 subfolders with the names of each of the neuropeptide families, as well as the results obtained with the IQ-TREE software. The folder named “<em>Trees</em>” contains the folder “<em>DNALM and kalign 2.0 default parameters</em>” containing the phylogenetic trees for each of the neuropeptide families, which were created with the Itol software. <strong>R5: BLAST alignment of the virtual peptide precursors</strong> Results of the BLAST alignment of the virtual peptides described by Martinez-Perez et al. (2007) with respect to the sequences in the GenBank database. The files follow the same nomenclature as in the folder named “<em>Carpeta 02 BLAST VP</em><strong>”</strong> in Repository 1. <strong>R6: Alignment of neuropeptide precursors</strong> “<em>DNALM and Kalign 2.0 default parameter</em>” folders. Each of these folders contains the alignments of the examined neuropeptide precursors from each family and each folder is named after the corresponding neuropeptide. The remaining files contain the alignments in ascending order in the evolutionary scale and are appropriately named after the corresponding neuropeptide. The file named “<em>All Sequence FASTA</em>” contains the sequences used in our study in FASTA format. <strong>R7: Phylogenetic clustering of the precursors </strong> “<em>DNALM and Kalign 2.0 default parameter</em>” folders. Both folders contain the phylogenetic tree clustering results from Supplementary Material 6, which were obtained using the DNA-LM y Kalign parameters and the IQ-TREE software. All analyses were conducted using the GUANE-1 supercomputer (Universidad Industrial de Santander). The phylogenetic clustering results of all of the precursors are contained in the folders with the respective precursor name. The folder also contains Figure 6, which was included in our main manuscript. Additionally, a folder entitled "Orthofinder and Robinson-Foulds" is included, which corresponds to the analyses carried out for: the Robinson-Foulds metric and the Orthofinder software.

# R1: 神经肽序列的构建与纯化 本研究使用10个关键词在GenBank数据库(GenBank)中检索LW、APGW、RPCH、脂肪动激素(AKH, Adipokinetic Hormone)、心激肽(CRZ, Corazonin)及促性腺激素释放激素(GnRH, Gonadotropin-Releasing Hormone)神经肽家族,关键词包括神经肽名称、前体缩写、前体全称、带“前原肽(prepropeptide)”字样的前体全称,以及上述术语的组合。通过GenBank数据库中的合适命令以FASTA格式(FASTA)下载候选序列。AKH神经肽家族根据文献公布的类群、氨基酸数量及序列进行分类。此外,通过BLAST比对(BLAST)在GenBank数据库中鉴定了ACP杂合家族。 ## C00: 神经肽前体(Neuropeptide Precursor) 创建8个以各神经肽家族首字母命名的文件夹,其中AKH家族文件夹为唯一包含4个子文件夹的文件夹。所有文件夹内均存储同类型文件:3个以神经肽首字母及所得结果命名的文本文件。带有“with codes(带编码)”字样的文件包含本研究生成的带编码序列,而带有“Full(完整)”字样的文档包含使用前述10个关键词检索GenBank数据库得到的结果,此类文件存储在名为“Fasta Keywords(FASTA关键词)”的文件夹中,每个文件对应单个关键词的检索结果。带有“selected EA(筛选进化分析)”字样的文件包含用于进化分析的筛选序列。 ## C01: BLAST ACP 名为“00 BLAST ACP”的文本文件包含从NCBI数据库获取的、以Callinectes toxotes转录组中的脂肪动激素/心激肽相关肽(Adipokinetic Hormone/Corazonin-related peptide)为检索序列的BLAST比对结果。名为“01 ACP Selected”的文件包含本研究选取的前体。所有序列均为FASTA格式(FASTA),并包含补充材料3“Database Sequences(数据库序列)”中汇总的编码。名为“02 ACP selected EA”的文件包含其他物种的ACP前体,用于Callinectes toxotes ACP的进化分析。名为“03 ACP ProP 1.0 Serv”的PDF文件包含使用前述软件对“02 ACP selected EA”文件中前体的蛋白酶切割位点分析结果。 ## C02: BLAST VP 该文件夹包含以Martinez-Perez等人2007年报道的虚拟肽序列为检索序列、针对NCBI数据库的BLAST比对结果。文件夹内包含7个文本文件,文件名对应所鉴定的前体及物种名称。此外,名为“Virtual peptides ProP 1.0 Serv(虚拟肽ProP 1.0 服务端)”的PDF文档包含使用前述软件生成的蛋白酶切割位点分析结果。 ## C03: 软件序列调试(Debugging sequences with software) 该文件夹包含3个子文件夹,存储本研究使用各软件通过合适关键词检测各神经肽序列所得的结果。名为“BioDataToolKit”的文件夹包含6个子文件夹,分别以各神经肽的缩写命名。此外,该文件夹内包含一个存储从GenBank数据库下载序列的文件,以及一个包含该软件生成详情的Microsoft Excel表格(Microsoft Excel),每个文件名对应各检索使用的关键词。本研究使用的软件可在以下仓库获取:https://github.com/rduarte24/BiodataToolkit。名为“Pro1.0Server”的文件夹组织方式与“BioDataToolKit”文件夹一致,按各神经肽家族划分,但每个神经肽文件夹内包含一个存储相关序列的文件,以及另一个存储通过该软件获得的神经肽前体内蛋白酶切割位点的文件。名为“Proteios”的文件夹包含7个文件,文件名对应使用软件分析的前体及以FASTA格式(FASTA)鉴定的序列。Proteios软件可在以下网址获取:https://github.com/Martin-Munive/Proteios。 ## C04: 用于进化分析的神经肽前体(Neuropeptide precursors for evolutionary analysis) 包含用于生成补充材料4和7中系统发育树的神经肽前体序列文件,每个文件名对应所分析的各神经肽名称。 ## R2: 转录组BLAST比对(Transcriptome BLAST) 包含使用Callinectes toxotes的AKH/CRZ相关肽(ACP)及C. arcuatus的心激肽(CRZ, Corazonin)序列进行BLAST比对所得结果的Microsoft Excel表格(Microsoft Excel)。在名为“C. toxotes”和“C. arcuatus”的工作表中汇总了以下信息:A列为神经肽名称,B列为物种名称,C至G列为BLAST比对结果,H列为GenBank蛋白质登录号,I列为前体序列。 ## R3: 神经肽数据库构建(Construction of neuropeptide database) 包含该数据库相关信息及本研究中分析的各神经肽前体详细描述的Microsoft Excel表格(Microsoft Excel)。该Excel文件包含7个工作表: 1. **神经肽(Neuropeptides)**:A列为降序编号的序列编号,B列为神经肽名称,C列为本研究使用的识别编码,D列为登录号,E至G列为物种分类学信息,H至L列为GenBank序列描述,M至N列为文献引用及链接。 2. **分类学(Taxonomy)**:从NCBI数据库获取的各受试物种的分类学描述。 3. **序列进化分析(Sequences evolutionary anal)**:C列为本研究开发的代码,D列为各神经肽的GenBank登录编码,E至F列为物种分类学详情。 4. **差异表(Table of differences)**:B列为相同序列的编码,C列为本研究选取的序列编码。 5. **已删除编码(Codes deleted)**:该工作表包含物种的登录编码及物种名称,但未包含神经肽前体的性质详情。 6. **论文序列(Sequences Paper)**:此前已在文献中报道、后续又在GenBank数据库中公布的神经肽序列。带星号标记的序列此前未在公共数据库中报道。同时包含本研究中用于标记序列的编码。 7. **关键词(Keywords)**:用于检索GenBank数据库以获取各神经肽家族成员的关键词。 ## R4: 计算机体外(In silico)验证、比对及系统发育关系 使用IQ-TREE软件(IQ-TREE),以DNA-LM和Kalign参数对各神经肽家族进行单次运行分析所得的系统发育树及结果。名为“RUN(运行)”的文件夹包含“DNALM and kalign 2.0 default parameters(DNALM及kalign 2.0默认参数)”子文件夹,两个文件夹均包含11个以各神经肽家族命名的子文件夹,以及使用IQ-TREE软件得到的结果。名为“Trees(树)”的文件夹包含“DNALM and kalign 2.0 default parameters(DNALM及kalign 2.0默认参数)”子文件夹,其中存放使用Itol软件构建的各神经肽家族的系统发育树。 ## R5: 虚拟肽前体的BLAST比对(BLAST alignment of the virtual peptide precursors) Martinez-Perez等人2007年描述的虚拟肽与GenBank数据库序列的BLAST比对结果。文件命名规则与仓库1中的“Carpeta 02 BLAST VP”文件夹一致。 ## R6: 神经肽前体比对(Alignment of neuropeptide precursors) 包含“DNALM and Kalign 2.0 default parameter(DNALM及Kalign 2.0默认参数)”文件夹,每个此类文件夹包含各家族受试神经肽前体的比对结果,每个文件夹以对应神经肽命名。其余文件包含进化尺度上按升序排列的比对结果,并以对应神经肽命名。名为“All Sequence FASTA(所有序列FASTA)”的文件包含本研究中使用的FASTA格式(FASTA)序列。 ## R7: 前体的系统发育聚类(Phylogenetic clustering of the precursors) 包含“DNALM and Kalign 2.0 default parameter(DNALM及Kalign 2.0默认参数)”文件夹,两个文件夹均包含补充材料6中的系统发育树聚类结果,该结果通过DNA-LM、Kalign参数及IQ-TREE软件获得。所有分析均在GUANE-1超级计算机(Universidad Industrial de Santander)上完成。所有前体的系统发育聚类结果存储在以对应前体名称命名的文件夹中,该文件夹还包含本研究主稿件中收录的图6。此外,还包含一个名为“Orthofinder and Robinson-Foulds(Orthofinder及Robinson-Foulds)”的文件夹,对应Robinson-Foulds度量及Orthofinder软件的分析结果。

提供机构:
Zenodo
创建时间:
2022-12-22
二维码
社区交流群
二维码
科研交流群
商业服务