遇见数据集

metagRoot: A comprehensive database of protein families associated with plant root microbiomes

收藏
Zenodo2025-06-13 更新2026-05-26 收录
官方服务:

资源简介:

metagRoot: A comprehensive database of protein families associated with plant root microbiomes AbstractThe plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71,091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, Hidden Markov Models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://www.metagroot.org --- Contents Summary.tsv.gz: A tab-delimited (tsv) table file containing summary information for each family: unique identifier for each family (e.g., F000292, F000581), the classification type of the family, the number of members within the family and the count of datasets and scaffolds associated with the family. Metadata.tsv.gz: A tab-delimited table file with the family metadata. Each row represents a family and includes details such as the number and percentage of metagenomes, metatranscriptomes, and isolates. Additionally, it provides information on the distribution of these families across different environments, including aerial roots, bulbs, endospheres, nodules, rhizoplanes, rhizospheres, stems, stem tubers, and unclassified regions. Domains.tsv.gz: PFAM domain annotations for each family. Each row includes information about the PFAM hit, the start and end positions of the hidden Markov model (HMM) alignment, and the corresponding genomic start and end positions. Additionally, it provides an accuracy score for the alignment. Sequences.tsv.gz: The representative sequences of the families. Each row includes the representative sequence length, the average length of sequences in that family, the header information and the sequence itself. FamFasta.tar.gz: TAR archive containing unaligned FASTA files for all protein families. Each file corresponds to a specific family and includes its protein sequences in standard FASTA format. These sequences can be used for downstream analyses such as alignment, annotation, and phylogenetic reconstruction. FamFastaAligned.tar.gz: TAR archive containing aligned and filtered sequences for all protein families, where each file includes the multiple sequence alignment of the representative protein sequences within a family. HMM.tar.gz: TAR archive containing HMM profile files for all protein families in the HMMER3 format, representing probabilistic models built from the aligned sequences of each family for use in sequence similarity searches and annotation. PDB.tar.gz: TAR archive containing 3D protein model predictions for each family, calculated with AlphaFold2/ColabFold

提供机构:
Zenodo
创建时间:
2025-06-13
二维码
社区交流群
二维码
科研交流群
商业服务