CARE
收藏资源简介:
CARE是由加州理工学院创建的一个用于酶分类和检索的基准数据集。该数据集包含高质量的酶和反应数据,以及它们关联的酶委员会(EC)编号,旨在通过机器学习方法评估和预测酶的功能。CARE数据集通过两个主要任务来评估模型:酶序列的分类和基于反应的酶检索。这些任务模拟了实际应用中的挑战,如未见过的蛋白质序列和未注释的反应。CARE数据集的创建过程涉及从多个数据库中提取和整理数据,确保数据的质量和可用性。该数据集的应用领域广泛,包括生物修复、塑料降解、基因编辑和药物合成等,旨在解决酶功能预测和设计中的关键问题。
CARE is a benchmark dataset for enzyme classification and retrieval developed by the California Institute of Technology. It contains high-quality enzyme and reaction data alongside their associated Enzyme Commission (EC) numbers, with the goal of evaluating and predicting enzyme functions via machine learning methods. The CARE benchmark assesses models through two core tasks: enzyme sequence classification and reaction-based enzyme retrieval. These tasks simulate real-world challenges such as unseen protein sequences and unannotated reactions. The creation of the CARE dataset involves extracting and curating data from multiple databases to ensure data quality and usability. This dataset has broad applications across areas including bioremediation, plastic degradation, gene editing, and pharmaceutical synthesis, and is designed to address critical challenges in enzyme function prediction and design.

- 1CARE: a Benchmark Suite for the Classification and Retrieval of Enzymes加州理工学院 · 2024年



