遇见数据集

VAPOR benchmark data, pretrained reward models, and ESM-2 logits cache for protein sequence optimization

收藏
Zenodo2026-05-15 更新2026-06-05 收录
官方服务:

资源简介:

This dataset provides the benchmark data and pretrained assets used byVAPOR, a Python toolkit for Pareto-leading protein sequence optimizationwithout invoking a protein language model during search. The deposit contains three archives: 1. vapor_benchmark_data.tar.gz: Benchmark DMS datasets used for protein sequence optimization experiments, including core benchmarks and prepared ProteinGym landscapes. 2. vapor_pretrained_models.tar.gz: Pretrained reward model ensemble checkpoints used by VAPOR. 3. vapor_esm2_logits_cache.tar.gz: Cached ESM-2 per-position token logits used for offline PLM distillation. These caches avoid runtime PLM calls during sequence optimization. These files are intended to be extracted into the VAPOR repository rootor release package directory before running the command-line interface.The source code is available separately as part of the VAPOR softwarepackage.

提供机构:
Zenodo
创建时间:
2026-05-15
二维码
社区交流群
二维码
科研交流群
商业服务