Bioinformatic Pipeline: Genomic Diversity Landscape Of The Honey Bee Gut Microbiota
收藏资源简介:
This data-set describes the full bioinformatic pipeline used to analyze 54 metagenomic samples of the honey bee gut microbiota. Each sample was isolated from an individual honey bee, and all samples originate from two colonies of the Engel laboratory at the University of Lausanne, Switzerland. The full raw data-set is available from the sequence-read archive: SRP150166. A publication based on this analysis is currently under review, with the title: "Genomic diversity landscape of the honey bee gut microbiota", and an upload to Biorxiv is also underway. The data-set contains tar-balls for the different main workflows of the analysis. Dowload and unpack to view the contents (tar -zxvf filename.tar.gz). For each workflow, all directories contain README.txt files, describing the contents of the directory. Due to size constraints, some intermediate files have been omitted, and some workflows are demonstrated for a subset of the data. However, the full analysis can be reproduced from the raw data, using the provided scripts. Scripts are included within workflow directories, and are also provided as a separate tar-ball for convenience. All perl-scripts come with documentation, which can be viewed by typing: "perl script_name.pl -h". For R scripts, the usage is indicated as a comment in the top lines of each script. Note that many of the scripts require specific input-files to be present in the run-directory. Their usage is demonstrated within the workflow directories in bash-scripts (*.sh). Commands used for generating plots and some statistics are given within workflow directories in text-files "R.commands" when applicable. Aside from custom code, the pipeline also utilizes various open-source Software packages, which are detailed in the file "software_dependencies.txt". Note, while many of the scripts will run fast on any computer, some steps of the pipeline are computationally demanding, and will require significant computing time, as well as storage space. When scripts are known to be time-consuming, this is indicated in the script help message.
本数据集详述了用于分析54份蜜蜂肠道菌群宏基因组样本的完整生物信息学分析流程。所有样本均取自单只蜜蜂,且全部来源于瑞士洛桑大学恩格尔实验室的两个蜂群。完整原始数据集可从序列读取归档库(Sequence Read Archive, SRA)获取,编号为SRP150166。 基于本分析的研究论文目前处于审稿阶段,论文标题为《蜜蜂肠道菌群的基因组多样性图谱》,同时该论文也正在上传至预印本平台bioRxiv。 本数据集包含对应分析各主要流程的tar压缩包。可通过命令`tar -zxvf filename.tar.gz`下载并解压以查看包内内容。每个流程对应的目录中均包含README.txt文件,用于说明该目录下的文件内容。受限于存储空间,部分中间文件已被省略,且部分流程仅以部分样本作为演示数据集。不过,借助本数据集提供的脚本,可基于原始完整数据复现全部分析流程。 分析脚本既存储于各流程目录中,同时也单独打包为tar压缩包以方便获取。所有Perl脚本(Perl script)均附带使用文档,可通过执行命令"perl script_name.pl -h"查看。R脚本(R script)的使用说明则以注释形式写在每个脚本的首行。请注意,多数脚本需要运行目录中存在指定的输入文件,其具体使用方法可在各流程目录下的Bash脚本(*.sh)中查看。如需生成图表与部分统计结果,相关命令会存储于各流程目录下的`R.commands`文本文件中(若适用)。 除自定义代码外,本分析流程还使用了多款开源软件包,相关详细信息可在`software_dependencies.txt`文件中查看。请注意,尽管多数脚本可在任意计算机上快速运行,但流程中的部分步骤计算量较大,需要耗费大量计算时间与存储空间。若脚本运行耗时较长,相关说明会在该脚本的帮助信息中注明。



