JMPhage Full Test Dataset
收藏资源简介:
Test dataset and reference outputs for the JMPhage pipeline (v1.0) This dataset provides a curated set of small input files and selected reference outputs for JMPhage v1.0, a modular one-stop pipeline for annotation, characterization, and comparative analysis of tailed phages (Caudoviricetes). It is intended for two purposes: To let new users verify that their local installation reproduces the expected behavior of each JMPhage subcommand. To serve as a lightweight worked example of what typical JMPhage outputs look like in both single-phage and multi-phage (joint) modes. Note on file selection: A complete JMPhage output directory can reach several gigabytes per phage once all intermediate files (BAM alignments, Pharokka native outputs, vConTACT2 intermediates, VIRIDIC working directories, GLUVAB per-tree scratch files, etc.) are retained. To keep this archive at a manageable size, only key end-user output files from each module are included. Intermediate files, native tool outputs, and per-tree scratch data have been removed. Users who run JMPhage locally on the provided inputs will obtain a superset of these outputs. Directory Structure and Dataset Contents single_phage/ — Three worked examples of single-genome analysis: vB_EcoM-Sa45lw_all_mode/: Escherichia coli phage vB_EcoM-Sa45lw analyzed in all mode (assembly + paired-end reads). z90_all_mode/: Aeromonas hydrophila jumbo phage Z90 (234 kb) analyzed in all mode (assembly + paired-end reads). vB_EcoS-UDS3lw_nomapping_mode/: Escherichia coli phage vB_EcoS-UDS3lw analyzed in no_mapping mode (assembly only). joint_phage/ — Two worked examples of multi-phage joint analysis: joint_all_mode/: Two phages (vB_EcoM-Sa45lw and Z90) analyzed jointly in all mode. Contains per-phage subdirectories plus a shared summary/ directory with the joint vConTACT2 network, joint ANI heatmap, and joint phylogenetic tree. joint_nomapping_mode/: Four Staphylococcus epidermidis phages (vB_Sep_steph1–4) analyzed jointly in no_mapping mode, illustrating the joint workflow when no paired-end reads are available. Files Retained per Module The archive keeps the following representative files per phage output directory: 1.read_mapping/: base_depth.tsv, depth_plot.pdf, and the phageterm/ report subdirectory 2.ORF_prediction/: phanotate.faa, <phage>.gbk, <phage>.gff 3.function_annotation/: <phage>_top_vog.tsv 4.genome_visualization/: genome_features_classified.tsv, protein_function_summary.tsv, genome_plot.pdf 5.shared_network/: c1.ntw 6.collinearity_analysis/: gene_functions.csv, <phage>.html 7.ANI_analysis/: Heatmap.PDF, sim_MA_genCol.csv 8.phylogenetic_tree/: <phage>_tree.pdf, small_tree/, big_tree/ For joint runs, the summary/ directory additionally contains: pipeline_report.tsv, shared_network/, ANI_analysis/, and phylogenetic_tree/. Omitted Files: Files omitted to save space include sorted BAM alignments, Pharokka native output subdirectories, raw DIAMOND/HMMER search tables, vConTACT2 intermediate files, VIRIDIC working directories, and GLUVAB per-tree scratch files. Related Resources JMPhage Source Code: https://github.com/linxiaoxu02/JMPhage.git JMPhage Reference Database (jm_db): https://doi.org/10.5281/zenodo.20838733 License Distributed under the GNU General Public License v3.0, consistent with the JMPhage source code.



