Phylogenetic distribution of the PHA metabolism in the archaeal domain
收藏资源简介:
This dataset contains the processed results of BLASTP searches against bacterial protein databases, parsed and annotated for downstream taxonomic analysis. XML output files from NCBI BLAST were converted to structured Excel format using a custom Python script. Each BLAST hit was filtered based on bit score (>100) and percent identity (>30%), and annotated with scientific name, accession, bit score, percent identity, alignment length, and E-value. Taxonomic information (phylum, class, and order) was automatically retrieved from the NCBI Taxonomy database via the Entrez API and appended to each entry. The final Excel file includes multiple sheets: Blast Hits – complete filtered BLAST dataset with taxonomy annotations. Summary_Phylum, Summary_Class, Summary_Order – frequency summaries across all hits. UniqueSpecies_Summary_Phylum, UniqueSpecies_Summary_Class, UniqueSpecies_Summary_Order – summaries calculated once per unique species to avoid redundancy. All scripts were written in Python (using xml.etree.ElementTree, pandas, Bio.Entrez, and openpyxl) and executed with an NCBI-friendly query delay to ensure compliance with usage policies. Author: Brendan Schroyen (Vrije Universiteit Brussel)



