pyTax4Fun2 example input files dataset
收藏资源简介:
General InformationThis is the example input files of pyTax4Fun2, a Python implementation of the Tax4Fun2 pipeline for prediction of habitat-specific functional profiles and functional redundancy based on 16S rRNA gene sequences. This dataset was generated from Zhou et al. [1] (NCBI Bioproject Accession: PRJNA1079022). Note that, not all of the original dataset used for building this dataset. In this dataset we only sample 3 experimental conditions: CK (Control bulk soil) SML (Bulk soil for low Cd-accumulating willow) SMH (Bulk soil for high Cd-accumulating willow) More details of the data that we use can be checked at the SraRunTable.csv that included in this dataset. This dataset divided into: asv-* (real ASV data) otu-* (real OTU data) toy_otu-* (toy OTU data, only contain 10 selected OTU only) Short Protocol This dataset was generated using this protocol: The FASTQ files (not included in this dataset), was downloaded from NCBI SRA [2,3]. Quality control performed by Falco [4] version 1.2.4 and MultiQC [5] version 1.33. Cutadapt [6] version 5.2 was used to clean the data from the custom primers (forward, CS1_515F: 5'-ACACTGACGACATGGTTCTACAGTGCCAGCMGCCGCGGTAA-3'; reverse, CS2_806R: 5'-TACGGTAGCAGAGACTTGGTCTGGACTACHVGGGTATCTAAT-3') used for the sequencing. Bowtie2 [7] version 2.5.4 was used to clean any possible of host (Salix jiangsunensis) fragment sequences. Since no genomic reference Salix jiangsunensis available at the time of our dataset generation, we use the genomic reference of its close relative, Salix arbutifolia (NCBI Genome Accession: GCA_025169955.1). Cleaned sequences then imported into QZA artifacts and processed using QIIME2 pipeline [8] version 2026.1. Under QIIME2 pipeline, denoising performed using DADA2 [9] to obtain amplicon sequence variants (ASVs) representative sequences and table. ASVs data later clustered into operational taxonomic units (OTUs) representative sequences and table using VSEARCH [10] using "cluster-features-open-reference" approach with this parameters set: --p-perc-identity 0.99 --p-strand both --p-threads 8. SILVA database [11] version 138.2 was used as reference sequence for feature clustering. Unless specifically noted, all programs are run under default parameters. References Zhou, J., Zhang, R., Wang, P., Gao, Y., Zhang, J. Responses of soil and rhizosphere microbial communities to Cd-hyperaccumulating willows and Cd contamination. BMC Plant Biology 24, 398 (2024). doi: 10.1186/s12870-024-05118-0. Katz, K., Shutov, O., Lapoint, R., Kimelman, M., Brister, J.R., O'Sullivan, C. The Sequence Read Archive: A decade more of explosive growth. Nucleic Acids Research 50, D387-D390 (2022). doi: 10.1093/nar/gkab1053. Sayers, E.W., Beck, J., Bolton, E.E., Brister, J.R., Chan, J., Connor, R., Feldgarden, M., Fine, A.M., Funk, K., Hoffman, J., Kannan, S., Kelly, C., Klimke, W., Kim, S., Lathrop, S., Marchler-Bauer, A., Murphy, T.D., O’Sullivan, C., Schmieder, C., Skripchenko, Y., Stine, A., Thibaud-Nissen, F., Wang, J., Ye, J., Zellers, E., Schneider, V.A, Pruitt, K.D. Database resources of the National Center for Biotechnology Information in 2025, Nucleic Acids Research 53, D20-D29 (2025). doi: 10.1093/nar/gkae979. de Sena Brandine, G., Smith, A.D. Falco: high-speed FastQC emulation for quality control of sequencing data. F1000Research 8, 1874 (2021). doi: 10.12688/f1000research.21142.2. Ewels, P., Magnusson, M., Lundin, S., Käller, M. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics 32, 3047-3048 (2016). doi: 10.1093/bioinformatics/btw354. Martin, M. Cutadapt Removes Adapter Sequences from High-Throughput Sequencing Reads. EMBnet.journal 17, 10-12 (2011). doi: 10.14806/ej.17.1.200. Langmead, B., Salzberg, S.L. Fast gapped-read alignment with Bowtie 2. Nature Methods 9, 357-359 (2012). doi: 10.1038/nmeth.1923. Bolyen, E. et al. Reproducible, interactive, scalable and extensible microbiome data science using QIIME 2. Nature Biotechnology 37, 852-857 (2019). doi: 10.1038/s41587-019-0209-9. Callahan, B., McMurdie, P., Rosen, M. Han, A.W., Johnson, A.J.A, Holmes, S.P. DADA2: High-resolution sample inference from Illumina amplicon data. Nature Methods 13, 581-583 (2016). doi: 10.1038/nmeth.3869. Rognes, T., Flouri, T., Nichols, B., Quince, C., Mahé, F. VSEARCH: a versatile open source tool for metagenomics. PeerJ 4, e2584 (2016). doi: 10.7717/peerj.2584. Chuvochina, M., Gerken, J., Frentrup, M., Sandikci, Y., Goldmann, R., Freese, H.M., Göker, M., Sikorski, J., Yarza, P., Quast, C., Peplies, J., Glöckner, F.O., Reimer, L.C. SILVA in 2026: a global core biodata resource for rRNA within the DSMZ digital diversity. Nucleic Acids Research 54, D334-D341 (2026). 10.1093/nar/gkaf1247.



