BGC-TFBS: predicted transcription factor binding motifs and regulatory annotations for biosynthetic gene cluster families
收藏资源简介:
Predicted transcription-factor binding motifs and regulatory annotations for the 1,263 BiG-FAM biosynthetic gene cluster (BGC) families with ≥100 members (760,746 BGCs, 17.5M protein-coding genes). Conserved upstream motifs were discovered per orthologous gene group (COG) with MEME across six settings (anr/zoops × max width 10/20/30); the zoops_maxw30 setting is the primary (showcase) set. Contents: gene tables with COG and TF-family annotation and protein sequences (all_genes.tsv.gz); BGC metadata with the mapping to original BiG-FAM accessions (bgc_metadata.tsv.gz); motif summaries for all settings (all_motifs.tsv.gz) and position weight matrices in MEME format per setting (motifs_*.meme, directly usable with FIMO/TOMTOM/MAST); per-gene motif occurrences with coordinates and p-values per setting (binding_sites_*.tsv.gz); regulator-to-motif pairing tables (tf_motif_pairings_*.tsv.gz); DNA-binding regulator annotations across all of BiG-FAM with protein sequences (tf_annotations.tsv.gz, tf_proteins.faa.gz); sequence similarity network packages for 20 transcription factor families (*_ssn_package.tar.gz) with PRODORIC-based validation; the PRODORIC reference set (please cite PRODORIC when using it); and the complete SQLite database behind the web resource (website.db.gz). README.md documents every file, column, and identifier convention; manifest.tsv lists md5 checksums. Enzyme families misannotated as regulators (e.g. Cytochrome P450) are excluded from all TF annotations.



