遇见数据集

Unifying Antimicrobial Peptide Datasets for Robust Deep Learning-Based Classification

收藏
DataCite Commons2025-05-16 更新2025-04-16 收录
官方服务:

资源简介:

Leguminous crops are vital to sustainable agriculture due to their ability to fix atmospheric nitrogen, improving soil fertility and reducing the need for synthetic fertilizers. Additionally, they are an excellent source of protein for both human consumption and animal feed. AntiMicrobial Peptides (AMPs), found in various leguminous seeds, exhibit broad-spectrum antimicrobial activity through diverse mechanisms, including interaction with microbial cell membranes and interference with cellular processes, making them valuable for enhancing crop resilience and food safety. In the field of plant sciences, computational biology methods have been instrumental in the discovery and optimization of AMPs. These methods enable rapid exploration of sequence space and the prediction of AMPs using deep learning technologies. Optimizing AMP annotations through computational design offers a strategic approach to enhance efficacy and minimize potential side effects, providing a viable alternative to conventional antimicrobial agents. However, the presence of overlapping sequences across multiple databases poses a challenge for creating a reliable dataset for AMP prediction. To address this, we conducted a comprehensive analysis of sequence redundancy across various AMP databases. These databases encompass a wide range of AMPs from different sources and with specific functions, including both naturally occurring and artificially synthesized AMPs. Our analysis revealed significant overlap, underscoring the need for a non-redundant AMP sequence database. We present the development of a new database that consolidates unique AMP sequences derived from leguminous seeds, aiming to create a more refined dataset for the binary classification and prediction of plant-derived AMPs. This database will support the advancement of sustainable agricultural practices by enhancing the use of plant-based AMPs in agroecology, contributing to improved crop protection and food security.

豆科作物对可持续农业至关重要,因其具备固定大气氮素的能力,可提升土壤肥力并降低合成肥料的使用需求。此外,它们还是人类膳食与畜禽饲料的优质蛋白质来源。从多种豆科种子中分离得到的抗微生物肽(AntiMicrobial Peptides, AMPs)可通过多种机制展现广谱抗菌活性,包括与微生物细胞膜相互作用、干扰细胞生理过程,因此在提升作物抗逆性与保障食品安全方面极具应用价值。在植物科学领域,计算生物学方法在抗微生物肽的发现与优化工作中发挥了关键作用。这类方法可快速探索序列空间,并借助深度学习技术实现抗微生物肽的预测。通过计算设计优化抗微生物肽的注释信息,是提升其功效、降低潜在副作用的有效策略,可为传统抗菌剂提供可行替代方案。然而,不同数据库间存在的重叠序列,为构建可靠的抗微生物肽预测数据集带来了挑战。为解决这一问题,我们对多款抗微生物肽数据库的序列冗余性开展了全面分析。这些数据库涵盖了不同来源、具备特定功能的各类抗微生物肽,包括天然存在与人工合成的抗微生物肽。分析结果显示存在显著的序列重叠问题,凸显了构建非冗余抗微生物肽序列数据库的必要性。我们在此介绍一款全新数据库的开发工作:该数据库整合了源自豆科种子的独特抗微生物肽序列,旨在为植物源抗微生物肽的二分类与预测任务构建更精准的数据集。该数据库将通过推动植物源抗微生物肽在农业生态学中的应用,助力可持续农业实践的发展,为强化作物保护与保障粮食安全提供支撑。

提供机构:
Recherche Data Gouv
创建时间:
2024-05-03
搜集汇总
背景与挑战
背景概述
该数据集整合了豆科植物种子中的抗菌肽序列,旨在解决现有AMP数据库中序列冗余问题,通过开发非冗余数据库支持深度学习分类,以促进可持续农业和食品安全的优化应用。数据集专注于提供独特AMP序列,用于稳健的二元分类和预测。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务