遇见数据集

PLMGuard

收藏
Zenodo2026-04-26 更新2026-05-26 收录
官方服务:

资源简介:

This dataset accompanies the PLMGuard repository and provides the data used to evaluate the reliability and interpretability of protein sequence search methods. Protein sequence search is a fundamental task in bioinformatics for identifying homologous proteins and inferring function. Traditional approaches such as BLASTp rely on sequence similarity, while recent methods leverage protein language models (PLMs) to capture deeper representations from sequence embeddings . However, embedding-based similarity can introduce opaque or misleading signals that are not always biologically meaningful. PLMGuard is a diagnostic framework for protein sequence search that probes whether similarity scores are biologically meaningful, semantically coherent, and resistant to manipulation. It helps distinguish trustworthy search signals from opaque embedding-based similarity across six complementary experiments. The dataset includes generated variant databases and modified versions of PLM-based search methods used in these experiments. It is intended for benchmarking and analyzing both traditional sequence-based methods and PLM-based retrieval approaches.

提供机构:
Zenodo
创建时间:
2026-04-26
二维码
社区交流群
二维码
科研交流群
商业服务