遇见数据集

February 26, 2026 (v1) Dataset Restricted Functional Assessment of Orphan Proteins in the Streptomyces Pan-Proteome through Genome-Wide Synteny Analysis - Additional script

收藏
Zenodo2026-03-26 更新2026-05-26 收录
官方服务:

资源简介:

This Zenodo folder concerns some scripts used for the paper "Functional Assessment of Orphan Proteins in the Streptomyces Pan-Proteome through Genome-Wide Synteny Analysis". Other information are aviable on DOI: 10.5281/zenodo.18784887 Analysis of Conserved Cluster in other Actinomycetota blast_clusters.py Given a list of clusters (input 1) and a multi-FASTA file (input 2), BLAST analyses were performed against multi-FASTA files contained within a specified directory. For each protein ID listed in the cluster dataset, the corresponding sequence was retrieved from the input multi-FASTA file and used as a query for similarity searches. filter_blast_results.py Filter the results of the previous script (coverage >70) Table_synteny Table with cluster name and syntenic proteins. Format: Cluster ID Syntenic proteinsCluster 1 KXXXXXCluster 1 KXXXXXCluster 1 KXXXXXCluster 2 KXXXXXCluster 2 KXXXXXCluster 3 KXXXXXCluster 3 KXXXXX... synteny_check.py Script to test synteny in other actinomycetes. Example: python synteny_check.py \ List_name_oranism.txt \ #List of clusters that included proteins identified in other actinomycetes name_oranism_KAAS.txt \ #protein ID with KEGG numbers (annotated genome) name_oranism_blast_results.tsv \ #results of BLAST Table_synteny.txt \ #Table with cluster name and syntenic proteins [synteny_results.tsv] Genomic position of proteins included in Conserved Clusters cluster_pivot.py Given a list of clusters (input 1) and a directory containing genomic coordinate files (input 2), the genomic position of each protein within each cluster was determined for all genomes analyzed.python cluster_pivot.py \ --clusters filtered_clusters.txt \ --coord_dir /Users/matteocalcagnile/Desktop/coordinates \ --output pivot.tsv \ pivot_to_binmatrix.p The dataset from cluster_pivot.py was subsequently used to quantify the number of clusters within defined genomic intervals (e.g., 50 kb), enabling downstream analysis of cluster distribution across genomes. python pivot_to_binmatrix.py \ --pivot pivot.tsv \ --output bin_matrix.tsv \ --bin_size 50000 CLUSER LIST FORMAT#Cluster 1WP_XXXWP_XXXWP_XXX…#Cluster 2WP_XXXWP_XXXWP_XXX…#Cluster 3WP_XXXWP_XXXWP_XXX…Genomic coordinate FORMATWP_XXX 509 2023WP_XXX 2091 2561WP_XXX 2558 3136WP_XXX 5482 6159WP_XXX... Enrichment analysis (FDR and OR calculation) Kegg_path KAAS annotation of genes syntenic to conserved clusters. Mapping file. population.txt Annotation performed with the KAAS tool of the reference genome: Streptomyces coelicolor A3(2). cluster_list List containing 2 columns: cluster names and KEGG number. Example format:Cluster ID Annotated genesCluster 1 KXXXXXCluster 1 KXXXXXCluster 1 KXXXXX... mapping.py Mapping script. Providing the cluster list and the mapping file produces a table with 3 columns: cluster name, KEGG number, and the KEGG pathway associated with the number. enrichment_fdr.pyFDR and OR calculation script. Provide the list produced by mapping.py and the reference genome file population.txt to calculate FDR and OR.

提供机构:
Zenodo
创建时间:
2026-03-25
二维码
社区交流群
二维码
科研交流群
商业服务