Gradients 1-3 polyA-selected transcripts per million, Gradients 3 depth profile polyA-selected processed metatranscriptomes
收藏资源简介:
G3_depth.assembledReads.id99.fasta.gz G3 depth profiles amino acid assembly Replicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinity Transcripts are reverse stranded but were assembled unstranded The assembly from each set of replicates was concatenated Each contig was six-frame translated Longest reading frame was selected for each contig (minimum amino acid length 100) Longest reading frame amino acid sequences were clustered at 99% amino acid identity The contigs in this file are the cluster representatives The code for processing and assemblying reads can be found here Amino acid format NPac.G3PA_depth.bf100.id99.nt.fasta.gz G3 depth profiles nucleotide assembly Replicate trimmed, quality-controlled reads in each direction (R1, R2) were combined across replicates before assemblying with trinity Transcripts are reverse stranded but were assembled unstranded The assembly from each set of replicates was concatenated Each contig was six-frame translated Longest reading frame was selected for each contig (minimum amino acid length 100) Longest reading frame amino acid sequences were clustered at 99% amino acid identity The contigs in this file are the nucleotide-encoded contigs that the amino acid-encoded cluster representatives are from The code for processing and assemblying reads can be found here Nucleotide format NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gz Taxonomic annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz) G3_depth.assembledReads.id99.fasta.gz was taxonomically annotated using diamond last common ancestor and MarFERReT Bash script used: diamond blastp --no-unlink -t ~/ -b 100 -c 1 -p 32 -d /mnt/nfs/projects/marferret/v1/data/marmicrodb/dmnd/MarFERReT.v1.1.MMDB.combined.dmnd -e 1e-5 --top 10 -f 102 -q G3_depth_assembledReads.id99.fasta -o NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab G3PA_depth.Pfam34.domtblout.tab.gz Pfam annotations for the G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz) G3_depth.assembledReads.id99.fasta.gz was functionally annotated against the Pfam database (version 34) using hmmsearch Bash script used: hmmsearch --cut_tc --domtblout G3PA_depth.Pfam34.domtblout.tab Pfam_34.0/Pfam-A.hmm G3_depth_assembledReads.id99.fasta G3PA_depth.raw.est_counts.csv.gz Estimated counts of transcripts from G3 depth profile samples mapped to G3 depth profile assembly (G3_depth.assembledReads.id99.fasta.gz) The code for processing, assemblying, and mapping reads can be found here Kallisto was used to map and outputs estimated counts (est_counts) The transcripts are reverse stranded but the G3 depth profiles were assembled unstranded so the reads were mapped against the assembly unstranded The G3 depth profile reads were mapped against the G3 depth profile assembly clustered at 99% amino acid identity in nucleotide space Bash script used: kallisto quant -i NPac.G3PA_depth.id99.nt.idx -o ${SAMPLE} --threads=${N_THREADS} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> G3PA_depth.${SAMPLE}.kallisto.log G2PA_incubations.est_counts.csv.gz Estimated counts of transcripts from G2 incubation samples mapped to G2 surface assembly Trimmed, quality-controlled reads from the G2 incubations were mapped to the G2 assembly ( Gradients2.MGL1704.PA.assemblies.tar.gz) The code for processing reads can be found here Kallisto was used to map and outputs estimated counts (est_counts) The transcripts are reverse stranded; the G2 surface samples were assembled reverse stranded; the G2 incubations were mapped reverse stranded against the G2 surface samples The G2 incubation reads were mapped against the G2 surface assembly clustered at 99% amino acid identity in nucleotide space Bash script used: kallisto quant --rf-stranded -i NPac.G2PA.bf100.id99.nt.idx -o ${SAMPLE} <(zcat ${LEFT_READS}) <(zcat ${RIGHT_READS}) >> ${SAMPLE}.kallisto.log At 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe added At 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe added At 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe added chlorophyllA_g2Incubations.csv Chlorophyll a measurements from the G2 incubations Samples for metatranscriptomes were collected after zero and 96 hours from the nutrient amendment experiments, along with onboard fluorometer measurements of chlorophyll a. Depth refers to the depth at which water was collected for the incubations. At 32.93 °N, LoNP: 0.5 uM NO3 and 0.05 uM PO4 added, HiNP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 0.5 nM Fe added At 37 °N, Fe: 1 nM Fe added, NP: 5 uM NO3 and 0.5 uM PO4 added, NPFe: 5 uM NO3 and 0.5 uM PO4 and 1 nM Fe added At 41.42 °N, LoFe: 0.3 nM Fe added, HiFe: 2 nM Fe added, NPFe: 10 uM NO3 and 1 uM PO4 and 2 nM Fe added The following datasets contain the transcripts per million per Pfam per taxonomic bin per sample: G1PA.tpm_counts.csv.gz G1 surface G2PA.tpm_counts.csv.gz G2 surface G3PA.tpm_counts.csv.gz G3 surface D1PA.tpm_counts.csv.gz Aloha diel G3PA_diel.tpm_counts.csv.gz G3 diel G2PA_incubations.tpm_counts.csv.gz G2 nutrient amendment incubations G3PA_depth.tpm_counts.csv.gz G3 depth profiles Transcripts per million were calculated with the following steps Divide the estimated number of reads mapped to each contig by its nucleotide length (in kilobases) to generate reads per kilobase (RPK). Sum the RPK by species and sample, then divide by one million to generate a conversion factor. Divide the RPK by the conversion factor to calculate transcripts per million per contig. Sum transcripts per million by Pfam for each species and sample. Transcript per million counts are only provided for annotated Pfams Transcripts per million were calculated using estimated counts outputted from kallisto Estimated count files for the G1-G3 surface, Aloha diel, and G3 diel samples can be found here Estimated count files for the G2 nutrient amendment incubation and G3 depth profile samples can be found in this repository: G2PA_incubations.est_counts.csv.gz and G3PA_depth.raw.est_counts.csv.gz respectively Transcripts per million were calculated per taxonomic bin Taxonomic annotations were generated using diamond last common ancestor and MarFERReT Taxonomic annotations for the G1-G3 surface, Aloha diel, and G3 diel samples can be found here Transcripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the taxonomic annotations for the G2 surface assembly can also found in the previous repository: NPac.G2PA.MarFERReT_v1.1_MMDB.lca.tab.gz Taxonomic annotations the G3 depth profile samples can be found in this repository: NPac.G3PA_depth.MarFERReT_v1.1_MMDB.lca.tab.gz Transcripts per million were calculated per Pfam Pfam annotations were generated using hmmsearch Pfam annotations for the G1-G3 surface, Aloha diel, and G3 diel samples were generated using version 35 of the Pfam database and can be found here Transcripts per million for the G2 incubations are based on trancripts from the G2 incubation samples mapped to the G2 surface assembly; the Pfam annotations for the G2 surface assembly were generated using version 35 of the Pfam database and can also found in the previous repository: G2PA.Pfam35.domtblout.tab.gz Pfam annotations the G3 depth profile samples were generated using version 34 of the Pfam database can be found in this repository: G3PA_depth.Pfam34.domtblout.tab.gz



