Synthyra/vector_embeddings
收藏资源简介:
该数据集名为Protify向量基准嵌入,是一个预计算的蛋白质嵌入数据集。它存储了多种蛋白质语言模型(如ESM2、ProtBert、ANKH等)和对照组(如Random、OneHot-Protein)的池化蛋白质嵌入,以gzip压缩的PyTorch文件格式(.pth.gz)提供。这些嵌入是序列级别的,采用均值和方差池化方法生成,旨在为Protify向量基准测试提供即用型嵌入,避免重复在本地GPU上计算相同基准序列的嵌入,从而加速下游任务,如经典机器学习、向量搜索、最近邻分析、低样本基准测试和模型比较。数据集包含约30个模型文件,总大小约189.22 GiB,用户可通过Hugging Face Hub下载单个或全部文件。
Precomputed pooled protein embeddings for the Protify vector benchmark. This dataset stores ready-to-use .pth.gz embedding artifacts for a broad panel of protein language models and controls. It is intended for fast downstream benchmarking in Protify without repeatedly embedding the same benchmark sequences on local GPUs, supporting tasks such as classical ML, vector search, nearest-neighbor analysis, low-shot benchmarking, or model comparison. The dataset includes approximately 30 model files with a total size of about 189.22 GiB, available for download via Hugging Face Hub.




