SAE Interpretability Dataset for Protein Language Models
收藏资源简介:
Processed dataset for residue-level interpretability of ESM-2 embeddings via Sparse Autoencoders (SAEs). Includes homology-clustered Swiss-Prot proteins with train/val/test splits (MMseqs2, 30% sequence identity), enriched UniProt feature annotations, and legacy annotation variants. Raw Swiss-Prot data (proteins.tsv, annotations.gff, swissprot_features_raw.tsv) should be downloaded directly from UniProt and is not included here. Files: clusters.tsv MMseqs2 clustering protein_splits.tsv The actual train/val/test split used proteins_with_split.tsv Proteins merged with split/cluster_id columns annotations_enriched.tsv Enriched annotations (annot_subtype, split, cluster_id). Main evaluation target
本数据集为基于稀疏自编码器(Sparse Autoencoders,SAEs)实现ESM-2嵌入残基级可解释性的预处理数据集。数据集包含经同源聚类的Swiss-Prot蛋白质序列,采用MMseqs2工具以30%序列同一性完成训练集/验证集/测试集划分;同时附带增强版UniProt特征注释与传统注释变体。原始Swiss-Prot数据(包括proteins.tsv、annotations.gff、swissprot_features_raw.tsv)需直接从UniProt官网下载,本数据集未包含此类原始文件。 数据集包含以下文件: 1. clusters.tsv:MMseqs2聚类结果文件 2. protein_splits.tsv:实际使用的训练/验证/测试集划分信息 3. proteins_with_split.tsv:合并了划分信息与聚类ID列的蛋白质数据 4. annotations_enriched.tsv:增强版注释文件(包含annot_subtype、split、cluster_id字段),为主要评估目标



