SAE Interpretability Dataset for Protein Language Models
收藏官方服务:
资源简介:
Processed dataset for residue-level interpretability of ESM-2 embeddings via Sparse Autoencoders (SAEs). Includes homology-clustered Swiss-Prot proteins with train/val/test splits (MMseqs2, 30% sequence identity), enriched UniProt feature annotations, and legacy annotation variants. Raw Swiss-Prot data (proteins.tsv, annotations.gff, swissprot_features_raw.tsv) should be downloaded directly from UniProt and is not included here. Files: clusters.tsv MMseqs2 clustering protein_splits.tsv The actual train/val/test split used proteins_with_split.tsv Proteins merged with split/cluster_id columns annotations_enriched.tsv Enriched annotations (annot_subtype, split, cluster_id). Main evaluation target
提供机构:
Zenodo创建时间:
2026-05-25



