遇见数据集

SAE Interpretability Dataset for Protein Language Models

收藏
Zenodo2026-05-25 更新2026-05-26 收录
官方服务:

资源简介:

Processed dataset for residue-level interpretability of ESM-2 embeddings via Sparse Autoencoders (SAEs). Includes homology-clustered Swiss-Prot proteins with train/val/test splits (MMseqs2, 30% sequence identity), enriched UniProt feature annotations, and legacy annotation variants. Raw Swiss-Prot data (proteins.tsv, annotations.gff, swissprot_features_raw.tsv) should be downloaded directly from UniProt and is not included here. Files: clusters.tsv MMseqs2 clustering protein_splits.tsv The actual train/val/test split used proteins_with_split.tsv Proteins merged with split/cluster_id columns annotations_enriched.tsv Enriched annotations (annot_subtype, split, cluster_id). Main evaluation target

提供机构:
Zenodo
创建时间:
2026-05-25
二维码
社区交流群
二维码
科研交流群
商业服务