Sigma70Pred: A highly accurate method for predicting sigma70 promoter in Escherichia coli K-12 strains
收藏资源简介:
Sigma70Pred: A Highly Accurate Method for Predicting Sigma70 Promoters in Escherichia coli K-12 Strains Sigma70Pred is a computational tool developed for predicting sigma70 promoters in Escherichia coli K-12 strains. Sigma70 factor plays a crucial role in prokaryotic transcription and regulates most housekeeping genes. Sigma70Pred uses nucleotide sequence-based features and machine learning models to classify DNA sequences as sigma70 promoters or non-promoters with high accuracy. Web Server: https://webs.iiitd.edu.in/raghava/sigma70pred/ Citation Patiyal, S., Singh, N., Ali, M. Z., Pundir, D. S., and Raghava, G. P. S. Sigma70Pred: A highly accurate method for predicting sigma70 promoter in Escherichia coli K-12 strains. Frontiers in Microbiology, 13, 1042127, 2022. https://doi.org/10.3389/fmicb.2022.1042127 About the Research Promoters are regulatory DNA regions located upstream of transcription start sites and are responsible for controlling gene expression. In prokaryotes, promoters are recognized by RNA polymerase together with sigma factors. Sigma70 is one of the most important sigma factors because it regulates the transcription of most housekeeping genes in Escherichia coli. Sigma70 promoters usually contain conserved sequence regions near the -10 and -35 positions upstream of the transcription start site. Accurate prediction of sigma70 promoters is important for understanding bacterial gene regulation, transcriptional control, genome annotation, and regulatory network analysis. Data Compilation: The benchmark dataset was obtained from RegulonDB 9.0 and contained 741 sigma70 promoters and 1400 non-promoters from Escherichia coli K-12. An independent dataset was created using RegulonDB 10.8 and contained 1134 sigma70 promoters and 638 non-promoters. Methodology: Sigma70Pred uses machine learning models trained on nucleotide sequence-based features. Around 8465 features were generated, including dinucleotide auto-correlation, dinucleotide cross-correlation, dinucleotide auto cross-correlation, Moran auto-correlation, normalized Moreau-Broto auto-correlation, pseudo tri-nucleotide composition, motif counts, GC skew, AT skew, and other nucleotide-based descriptors.



