遇见数据集

Test dataset for MacSyFinder (v2) and expected output files

收藏
Figshare2022-11-21 更新2026-04-28 收录
官方服务:

资源简介:

We present here a sample of a genomic dataset, to test MacSyFinder on it (the complete sequence of Vibrio cholerae O1 biovar El Tor str. N16961 chromosome I), along with the expected output files (on the form of compressed folders). The chromosome to annotate is presented as a multi-FASTA file of the proteins ordered as the genes encoding them. An annotation of the protein secretion systems and appendages was run on the genome, using the macsyfinder set of models ("macsy-model") TXSScan, V1.1.1 in the case of these examples. There are two output files offered, the one expected with the "ordered" genome mode of annotation, and the other with the "unordered" mode of genome annotation. The following command lines were used to obtain the output files: First, the genome is downloaded From the present page. Second, the MacSyFinder program is installed. Have a look here for procedure: https://github.com/gem-pasteur/macsyfinder Third, the TXSScan models for annotation of secretion systems are installed. The command line is the following: > macsydata install TXSScan # Install the latest version of TXSScan Finally, MacSyFinder is run on the genome, here using 8 workers for the HMM search ("-w 8" option): In "ordered" mode: > macsyfinder --sequence-db VICH001.B.00001.C001.fasta -o macsyfinder_TXSScan_VICH001_ordered --models TXSScan all --db-type ordered_replicon -w 8 # specified output folder: macsyfinder_TXSScan_VICH001_ordered In "unordered" mode: > macsyfinder --sequence-db VICH001.B.00001.C001.fasta -o macsyfinder_TXSScan_VICH001_unordered --models TXSScan all --db-type unordered -w 8 # specified output folder: macsyfinder_TXSScan_VICH001_unordered The documentation on the generated output files Can be consulted here.

本数据集提供一组基因组测试样本,用于在其上测试MacSyFinder工具,该样本为霍乱弧菌O1生物型El Tor菌株N16961的1号染色体完整序列,同时附带以压缩文件夹形式提供的预期输出文件。待注释的染色体以多FASTA(multi-FASTA)文件形式提供,文件内的蛋白序列按照其编码基因的顺序排列。本示例采用macsy-model模型集的TXSScan(版本V1.1.1)对该基因组开展蛋白分泌系统及附属结构的注释分析。本次提供两类输出文件,分别对应基因组注释的"有序(ordered)"模式与"无序(unordered)"模式的预期结果。以下为获取输出文件所用的命令行与步骤:首先,从本页面下载基因组序列;其次,安装MacSyFinder程序,安装流程可参考:https://github.com/gem-pasteur/macsyfinder;第三,安装用于分泌系统注释的TXSScan模型,命令如下:> macsydata install TXSScan # 安装最新版TXSScan;最后,在基因组上运行MacSyFinder,本次使用8个线程进行HMM(隐马尔可夫模型,Hidden Markov Model)搜索(参数"-w 8"):在"有序"模式下:> macsyfinder --sequence-db VICH001.B.00001.C001.fasta -o macsyfinder_TXSScan_VICH001_ordered --models TXSScan all --db-type ordered_replicon -w 8 # 指定输出文件夹:macsyfinder_TXSScan_VICH001_ordered;在"无序"模式下:> macsyfinder --sequence-db VICH001.B.00001.C001.fasta -o macsyfinder_TXSScan_VICH001_unordered --models TXSScan all --db-type unordered -w 8 # 指定输出文件夹:macsyfinder_TXSScan_VICH001_unordered;关于生成的输出文件的详细文档可在此处查阅。

创建时间:
2022-11-21
二维码
社区交流群
二维码
科研交流群
商业服务