BIOSYSMOdb: Curated Database for Biodegradation and Bioremediation
收藏资源简介:
BIOSYSMOdb is a comprehensive and integrative database developed as part of BIOSYSMO project. This resource centralizes data on metabolic pathways, reactions, enzymes, and degradative organisms to address soil contamination caused by industrial, agricultural, and urban activities. BIOSYSMOdb serves as a bridge between computational and experimental research, offering a unified platform to accelerate bioremediation solutions. Dataset Description BIOSYSMOdb integrates curated and synthesized data from major public repositories: EAWAG BBD, MibPOPdb, MetaCyc, Uniprot, and KEGG. The database includes: Chemical level: Details on compounds relevant for biodegradation. Metabolic level: Data on pathways, reactions, enzymes, and organisms associated with degradation. Organism level: Information on degradative organisms and their genomic data. Protein level: Information on enzymes in charge of each reaction and their sequence data associates (if available) Data Structure The following files are included in the dataset: BIOSYSMOdb_Compounds_chemical_iden_v1.0.csv: Compounds identifiers iferred for other databases BIOSYSMOdb_Compounds_chemical_info_v1.0.csv: Compounds information collected from public sources BIOSYSMOdb_Compounds_onthology_cod_v1.0.csv: Compounds onthology codes derived from Classyfire BIOSYSMOdb_Compounds_onthology_term_v1.0.csv: Compounds onthology terms derived from Classyfire BIOSYSMOdb_Pathways_v1.0.csv: Pathways dataset BIOSYSMOdb_Reactions_v1.0.csv: Reactions dataset (containing substrates, products, enzymes and pathways associated) BIOSYSMOdb_Enzymes_v1.0.csv: Reactions dataset (containing reactions associated) BIOSYSMOdb_Compounds_v1.0.csv: Compounds principal dataset BIOSYSMOdb_Organisms_v1.0.csv: Organisms principal dataset (containing pathways associated and NCBI Genome ID when available) CSV Descriptions Compound ID: Unique identifier for each compound. Pathway Name: Name of the metabolic pathway. Reaction ID: Identifier for individual reactions. Enzyme/Protein ID: Unique identifier for associated enzymes. Organism Name: Name of the degradative organism. Jupyter Notebook for querying BIOSYSMOdbTo facilitate data exploration and connections within the CSV files, a Jupyter Notebook, BIOSYSMO_database_queries, has been created. This notebook enables users to analyze relationships between different datasets and execute relevant queries efficiently. Data Sources & Licenses This database includes data derived from diverse databases: EAWAG BBD: Data on biodegradation of persistent organic pollutants. Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) MibPOPdb: Focused on microbial degradation of xenobiotics. Creative Commons Attribution 4.0 International (CC BY 4.0) license. MetaCyc: Comprehensive metabolic pathway database. Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) KEGG: Genomic integration and metabolic networks. Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) UniProt: Protein sequences. Creative Commons Attribution 4.0 International (CC BY 4.0) license. NCBI Genome: Organism Genomes. This database is public. Pubchem: Chemical Compounds. this database is public. CHebi: Chemical Compounds. Creative Commons Attribution 4.0 International (CC BY 4.0) license. Licensing and Attribution This dataset is shared under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0). Please credit BIOSYSMOdb and the original sources (EAWAG BBD, MibPOPdb, MetaCyc, ChEBI, Pubchem and KEGG) in any use or derivative works. BIOSYSMOdb was developed as part of the BIOSYSMO project, which has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No. 101060211. Acknowledgments- MetaCyc, KEGG, EAWAG BBD, UniProt, NCBI Genome, PubChem, ChEBI, and MibPOPdb – For providing essential data that supported the curation of BIOSYSMOdb.- BIOSYSMO consortium – For their contributions to the database’s design and development. We extend our gratitude to the Horizon Europe programme and the European Union for their support in advancing research on bioremediation and biodegradation. Contact For inquiries, please contact: - Contact Name: Main Researcher: Marta Franco de Benito, MsC or Project Coordinator: Sara Gil Guerrero, PhD - Email: marta.franco@idener.ai // sara.gil@idener.ai - Institution: IDENER.AI
BIOSYSMOdb是一款综合性整合数据库,作为BIOSYSMO项目的组成部分开发完成。该资源汇聚了代谢通路、反应、酶类以及降解生物相关数据,旨在应对工业、农业与城市活动引发的土壤污染问题。BIOSYSMOdb搭建起计算研究与实验研究之间的桥梁,提供统一平台以加速生物修复解决方案的研发进程。 数据集说明 BIOSYSMOdb整合了来自多个主流公共数据库的经人工审核整理与合成的数据,包括EAWAG BBD、MibPOPdb、MetaCyc、通用蛋白质知识库(UniProt)以及京都基因与基因组百科全书(KEGG)。该数据库涵盖以下四类内容: 1. 化学层级:与生物降解相关的化合物详细信息。 2. 代谢层级:与降解过程相关的通路、反应、酶类以及生物的关联数据。 3. 生物层级:降解生物的相关信息及其基因组数据。 4. 蛋白质层级:负责各反应的酶类信息及其关联的序列数据(若有提供)。 数据结构 本数据集包含以下文件: - BIOSYSMOdb_Compounds_chemical_iden_v1.0.csv:源自其他数据库的化合物标识符文件 - BIOSYSMOdb_Compounds_chemical_info_v1.0.csv:从公共来源收集的化合物信息文件 - BIOSYSMOdb_Compounds_ontology_cod_v1.0.csv:源自Classyfire的化合物本体编码文件 - BIOSYSMOdb_Compounds_ontology_term_v1.0.csv:源自Classyfire的化合物本体术语文件 - BIOSYSMOdb_Pathways_v1.0.csv:代谢通路数据集 - BIOSYSMOdb_Reactions_v1.0.csv:反应数据集(包含底物、产物、关联酶类与通路) - BIOSYSMOdb_Enzymes_v1.0.csv:酶数据集(包含关联的反应信息) - BIOSYSMOdb_Compounds_v1.0.csv:核心化合物数据集 - BIOSYSMOdb_Organisms_v1.0.csv:核心生物数据集(包含关联的代谢通路以及可用的NCBI基因组ID) CSV文件字段说明 各CSV文件的核心字段说明如下: - 化合物ID:每个化合物的唯一标识符。 - 通路名称:代谢通路的名称。 - 反应ID:单个反应的标识符。 - 酶/蛋白质ID:关联酶类的唯一标识符。 - 生物名称:降解生物的名称。 BIOSYSMOdb查询用Jupyter Notebook 为便于用户探索CSV文件内的数据并建立数据关联,我们开发了一款名为BIOSYSMO_database_queries的Jupyter Notebook。该工具可帮助用户分析不同数据集间的关联关系,并高效执行相关查询操作。 数据来源与许可协议 本数据库的数据源自多个不同的数据库: 1. EAWAG BBD:持久性有机污染物生物降解相关数据,采用知识共享署名-非商业性使用-相同方式共享4.0国际许可(Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International,CC BY-NC-SA 4.0) 2. MibPOPdb:聚焦于异生物质的微生物降解研究,采用知识共享署名4.0国际许可(Creative Commons Attribution 4.0 International,CC BY 4.0) 3. MetaCyc:综合性代谢通路数据库,采用CC BY-NC-SA 4.0许可 4. KEGG:基因组整合与代谢网络数据库,采用CC BY-NC-SA 4.0许可 5. UniProt:蛋白质序列数据库,采用CC BY 4.0许可 6. NCBI Genome:生物基因组数据库,为公共开放资源 7. PubChem:化合物数据库,为公共开放资源 8. ChEBI:化合物数据库,采用CC BY 4.0许可 许可与署名要求 本数据集采用CC BY-NC-SA 4.0许可协议进行共享。在任何使用或衍生作品中,请务必注明BIOSYSMOdb以及原始数据源(EAWAG BBD、MibPOPdb、MetaCyc、ChEBI、PubChem及KEGG)。 BIOSYSMOdb作为BIOSYSMO项目的一部分开发完成,该项目获得了欧盟地平线欧洲研究与创新计划的资助,资助协议编号为101060211。 致谢 - 感谢MetaCyc、KEGG、EAWAG BBD、UniProt、NCBI Genome、PubChem、ChEBI以及MibPOPdb提供的关键数据,为BIOSYSMOdb的整理工作提供了支撑。 - 感谢BIOSYSMO联盟为数据库的设计与开发所做出的贡献。 我们谨向地平线欧洲计划与欧盟致谢,感谢其对生物修复与生物降解研究的支持。 联系方式 如有任何疑问,请联系: - 联系人:主要研究员:Marta Franco de Benito,理学硕士;或项目协调员:Sara Gil Guerrero,哲学博士 - 邮箱:marta.franco@idener.ai // sara.gil@idener.ai - 所属机构:IDENER.AI



