Composite non-redundant metagenomic gene sequences from extreme environments of Pakistan
收藏资源简介:
This dataset contains composite non-redundant predicted gene (coding DNA) sequences derived from shotgun metagenomic analyses of environmental samples collected from extreme habitats in Pakistan, including geothermal hot springs, arid desert soil, and salt mine sediments. The dataset represents processed metagenomic outputs generated after quality control, assembly, gene prediction, and redundancy reduction. The primary analyses reported in the associated study are based on gene-centric functional profiling, comparative gene content analysis, and annotation-driven inference of metabolic potential. These analyses rely on the presence, absence, and annotation of non-redundant gene sequences rather than on read-level abundance estimation or re-assembly procedures. Accordingly, the non-redundant gene sequences provided here constitute the exact input data used for downstream functional annotation, pathway reconstruction, and comparative analyses. Each FASTA file contains a composite set of non-redundant genes corresponding to the same biological replicate number (R1, R2, or R3), pooled across all sampling sites. This structure reflects the analytical design of the study and enables direct reproduction of all reported gene-based results, including functional classification, enrichment analyses, and comparative assessments across replicates. Raw sequencing reads are not included, as they are not required to reproduce the gene-centric analyses, annotations, or conclusions presented in the study. The provided sequences represent the finalized and analysis-ready dataset upon which all results are based.



