遇见数据集

BioRemPP Database: A Curated Compound-Centric Resource for Bioremediation Potential Profiling

收藏
Zenodo2026-06-21 更新2026-06-28 收录
官方服务:

资源简介:

Overview The BioRemPP Database (Bioremediation Potential Profile Database) is a curated, integrated resource designed to support environmental bioremediation research by systematically linking chemical compounds, genes, KEGG Orthology groups, enzyme activities, biochemical reactions, degradation pathways, toxicity predictions, and regulatory frameworks. The database addresses a critical gap in bioremediation science: the absence of a unified and standardized resource connecting environmentally relevant pollutants with their potential biological transformation mechanisms across multiple biochemical, genomic, toxicological, and regulatory knowledge bases. Version 1.1.0 expands the original compound–gene integration model by incorporating Enzyme Commission numbers, KEGG reaction identifiers, and biochemical reaction descriptions. Scientific Rationale Environmental contamination by xenobiotic compounds—including chlorinated solvents, polyaromatic hydrocarbons, pesticides, petroleum-derived hydrocarbons, metals, and other hazardous substances—poses significant ecological and public-health challenges. Although substantial information exists regarding microbial biodegradation capabilities, this knowledge remains fragmented across biochemical databases, degradation-gene repositories, toxicological prediction platforms, and regulatory frameworks. BioRemPP systematically integrates these sources into a unified, FAIR-oriented framework that supports pollutant-centered functional exploration, comparative genomic profiling, pathway coverage analysis, organism prioritization, and hypothesis generation for bioremediation research. The associations available in BioRemPP are database-derived and should not be interpreted as direct experimental evidence of biodegradation activity or environmental remediation efficiency. Database Contents (v1.1.0) This release contains: 123,762 database records representing compound–KO–enzyme–reaction–regulatory agency combinations 384 unique chemical compounds with standardized identifiers and chemical classifications 1,543 KEGG Orthology identifiers for functional annotation 1,517 gene symbols 1,422 functional gene or enzyme descriptions 206 standardized enzyme-activity terms 994 Enzyme Commission numbers 1,939 KEGG reaction identifiers 12 chemical compound classes 9 regulatory agency or priority-list codes 370 compounds with ToxCSM predictions 31 toxicological endpoints for each compound covered by ToxCSM The 123,762 records constitute a denormalized analytical table. The number of records should not be interpreted as the number of unique compounds, independent biological observations, or experimentally validated degradation reactions. Data Sources and Integration Data were curated and integrated from multiple authoritative and specialized resources. Regulatory Frameworks ATSDR — Agency for Toxic Substances and Disease Registry Substance Priority List EPA — United States Environmental Protection Agency priority compound lists CONAMA — Brazilian National Environment Council regulations IARC — International Agency for Research on Cancer classifications, including Groups 1, 2A, and 2B European Union Water Framework Directive Priority Substances European Environmental Priority Chemicals Canadian Environmental Protection Act Priority Substances List Functional and Biochemical Annotations KEGG Compound — standardized compound identifiers and compound names KEGG Orthology — KO identifiers, gene symbols, and functional descriptions KEGG Enzyme and Reaction resources — EC numbers, reaction identifiers, and biochemical reaction descriptions KEGG xenobiotic degradation pathways — pollutant- and xenobiotic-associated pathway annotations HADEG — hydrocarbon aerobic degradation-specific gene, enzyme, and pathway coverage BlastKOALA and EggNOG-mapper — sequence-based functional annotation support Chemical Classification ChEBI — standardized chemical identifiers, molecular representations, and compound classification support Toxicological Annotations ToxCSM — machine-learning-based multi-endpoint toxicity predictions, including nuclear, stress-response, genomic, environmental, and dose-related endpoints Integration Framework and External Database Contributions BioRemPP is not intended to replace or reproduce the complete contents of existing databases. Instead, it functions as an integration layer that establishes compound-centered relationships across independent external resources, each contributing distinct and complementary information. The BioRemPP integration workflow standardizes identifiers and generates relational mappings that allow users to navigate among environmentally relevant compounds, KO groups, genes, enzyme activities, EC numbers, biochemical reactions, degradation pathways, toxicity predictions, and regulatory classifications. The primary version 1.1.0 source table contains 123,762 records, 384 unique compounds, and 1,543 unique KO identifiers. Complementary pathway and toxicity resources are connected through standardized compound and KO cross-references. The HADEG source dataset contains 71 pathway names across five compound-pathway categories, while the KEGG degradation source dataset contains 20 pathway names. Following integration, filtering, and deduplication, the BioRemPP Database Explorer runtime database contains 66 pathway–source records. The following external resources contribute specific information layers to BioRemPP: External Resource Contribution to BioRemPP Quantitative Coverage Relationship Type KEGG Compound and KEGG Orthology Compound identifiers, KO identifiers, gene symbols, gene descriptions, and functional annotations 384 compounds and 1,543 KO identifiers Cross-reference through compound and KO identifiers KEGG Enzyme and Reaction resources Enzyme Commission numbers, reaction identifiers, and biochemical reaction descriptions 994 EC numbers and 1,939 KEGG reaction identifiers Cross-reference through KO, EC, and reaction identifiers KEGG Degradation Pathways Xenobiotic and pollutant metabolism pathway associations 20 source pathway names Cross-reference through KO identifiers HADEG Degradation-specific gene and pathway coverage for hydrocarbons, polymers, alkanes, alkenes, aromatics, and biosurfactant-related processes 71 source pathway names across five compound-pathway categories Cross-reference through KO identifiers ChEBI Standardized chemical identifiers, molecular representations, and compound classification support Chemical normalization and classification support for the compound collection Cross-reference through chemical identifiers ToxCSM Machine-learning-based toxicity predictions as additional annotation layers 370 compounds with 31 toxicity endpoints, corresponding to 11,470 compound–endpoint records Cross-reference through compound identifiers Regulatory Frameworks Priority compound lists and hazard classifications defining environmentally relevant compounds 9 regulatory codes and 806 unique compound–agency relationships Compound inclusion and classification criteria What BioRemPP Adds Compound-centric relational structure — Links compounds to KO groups, genes, enzyme activities, EC numbers, biochemical reactions, pathways, toxicity predictions, and regulatory classifications through standardized identifiers. Cross-database harmonization — Resolves identifier inconsistencies and creates interoperable relationships among KEGG, ChEBI, HADEG, ToxCSM, and regulatory sources. Biochemical reaction-level resolution — Extends the original gene-centered model by incorporating EC numbers, KEGG reaction identifiers, and biochemical reaction descriptions. Regulatory context integration — Associates functional and biochemical annotations with environmental priority status from multiple international regulatory frameworks, structuring the database around environmentally relevant pollutants rather than pathway catalogs alone. Toxicological context integration — Connects compound biodegradation-related annotations with computational toxicity predictions across 31 endpoints. Analytical framework — Provides a structured 11-field data table containing 123,762 records, optimized for bioremediation potential profiling, functional coverage analysis, comparative genomics, and sample comparison. Documented data quality — Core identifier and annotation fields are complete across the dataset. EC-number coverage is 99.22%, while KEGG reaction identifier and reaction-description coverage is 98.20%. Users seeking complete primary records, original pathway diagrams, experimentally measured toxicity data, or detailed toxicological reports should consult the corresponding source databases directly. BioRemPP facilitates navigation by maintaining traceable identifiers and cross-references to the original resources. Reference Genomes For demonstration and validation purposes, the database includes functional annotations from nine representative RefSeq genomes spanning principal bioremediation-relevant groups: Bacteria: Acinetobacter baumannii, Enterobacter asburiae, Pseudomonas aeruginosa Fungi: Aspergillus nidulans, Fusarium graminearum, Cryptococcus gattii Microalgae/Cyanobacteria: Chlorella variabilis, Nannochloropsis gaditana, Synechocystis sp. These reference genomes support demonstration of functional profiling and comparison workflows and should not be interpreted as representing the complete taxonomic scope of the database. File Formats The primary BioRemPP v1.1.0 dataset is provided as a UTF-8 encoded, semicolon-delimited CSV file. Primary File biorempp_database_v1.1.0.csv File Properties Property Value Version 1.1.0 Records 123,762 Columns 11 Approximate file size 32 MB Character encoding UTF-8 Field delimiter Semicolon Text qualifier Double quotation mark Header row Yes Dataset Fields Field Description cpd KEGG Compound identifier compoundclass Standardized chemical compound class ko KEGG Orthology identifier ec Enzyme Commission number, when available reaction KEGG Reaction identifier, when available reaction_description Description of the associated biochemical reaction referenceAG Regulatory agency or priority-list code compoundname Standardized compound name genesymbol Gene symbol associated with the KO group genename Functional gene or enzyme description enzyme_activity Standardized enzyme-activity term Detailed field descriptions, controlled vocabularies, identifier formats, entity relationships, source provenance, missing-value interpretation, and data-quality notes are included in the accompanying BioRemPP documentation. Associated Web Server The BioRemPP web server provides interactive visualization and analysis tools for exploring the database through compound, gene, pathway, biochemical reaction, regulatory, and toxicological perspectives. The web server is available at: https://bioinfo.imd.ufrn.br/biorempp/ The platform includes eight analytical modules and 56 use cases supporting functional exploration, comparative profiling, data visualization, and hypothesis generation. The Zenodo release provides a versioned archival copy of the database to support reproducibility, citation, and independent reuse. License Original BioRemPP curation, integration structures, documentation, and project-generated materials are released under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Third-party database content, identifiers, annotations, and derived information remain subject to the licenses, access conditions, citation requirements, and terms of use established by their respective providers. The BioRemPP CC BY 4.0 license does not replace, override, sublicense, or modify the terms associated with KEGG, HADEG, ChEBI, ToxCSM, or regulatory source materials. Users requiring complete primary records, pathway diagrams, detailed toxicological reports, redistribution rights, or commercial reuse rights should consult the corresponding source provider and its current licensing terms. Version History v1.0.0 — December 2025: Initial release containing 10,869 records and 8 fields. v1.1.0 — April 2026: Expanded release containing 123,762 records and 11 fields. Added the ec, reaction, and reaction_description fields; expanded the integration of KEGG Orthology, enzyme, and reaction information; and updated the database using KEGG Release 117.0+. Contact For questions, corrections, or feedback, please contact: biorempp@gmail.com Technical issues may also be submitted through the project repository: https://github.com/BioRemPP/biorempp_web/issues Keywords bioremediation, biodegradation, xenobiotics, environmental microbiology, functional genomics, bioinformatics, environmental biotechnology, KEGG, KEGG Orthology, KEGG reactions, Enzyme Commission numbers, HADEG, ChEBI, ToxCSM, pollutants, regulatory compounds, compound–gene associations, compound–reaction associations, toxicology, metagenomics Metadata Fields Field Value Resource Type Dataset Title BioRemPP Database: A Curated Compound-Centric Resource for Bioremediation Potential Profiling Version 1.1.0 License Creative Commons Attribution 4.0 International for original BioRemPP contributions; third-party terms remain applicable Language English Subjects Environmental Sciences, Bioinformatics, Microbiology, Biotechnology, Toxicology, Computational Biology

提供机构:
Zenodo
创建时间:
2026-06-21
二维码
社区交流群
二维码
科研交流群
商业服务