遇见数据集

CAZy-HABench30K: A Homology-Aware Benchmark Dataset for CAZy Enzyme Family Classification

收藏
Mendeley Data2026-07-03 收录
官方服务:

资源简介:

This dataset contains the CAZy-HABench30K benchmark used for homology-aware CAZy enzyme family classification. The dataset includes 30,000 protein sequences collected from UniProt and annotated with CAZy family and class labels. It covers 60 CAZy families across the six major CAZy classes: glycoside hydrolases, glycosyltransferases, polysaccharide lyases, carbohydrate esterases, auxiliary activities, and carbohydrate-binding modules. The package provides FASTA files, family/class labels, MMseqs2 cluster assignments, and fixed train/validation/test splits generated using a cluster-disjoint homology-aware splitting protocol. The final split contains 21,000 training sequences, 4,500 validation sequences, and 4,500 test sequences. The dataset also includes prediction-level outputs, homology leakage analysis results, threshold sensitivity tables, ablation results, truncation analysis outputs, and per-family performance summaries used in the associated manuscript. This benchmark is intended to support reproducible evaluation of CAZy enzyme family classification models under homology-aware conditions. It is particularly suitable for studying performance inflation caused by random sequence-level splitting, remote-homology generalization within known CAZy families, protein language model evaluation, and comparison with homology-based inference methods such as MMseqs2 nearest-neighbor transfer. No external funding was received for the creation, preparation, or publication of this dataset. Keywords: CAZy; enzyme family classification; homology-aware benchmark; MMseqs2; protein language models; ESM-2; homology leakage; UniProt; bioinformatics; benchmark dataset

创建时间:
2026-06-23
二维码
社区交流群
二维码
科研交流群
商业服务