DART-Eval
收藏资源简介:
DART-Eval是由斯坦福大学开发的一个全面的DNA语言模型评估基准,专注于调控DNA。该数据集包含230万条调控DNA序列,旨在评估DNA语言模型在零样本、探针和微调设置下的性能。数据集的创建过程包括从ENCODE项目中精选的cis-调控元件(cCREs)和通过保持二核苷酸频率生成的合成负样本。DART-Eval的应用领域包括序列基序发现、细胞类型特异性调控活性预测和调控遗传变异的反事实预测,旨在解决调控DNA序列的复杂性问题。
DART-Eval is a comprehensive evaluation benchmark for DNA language models developed by Stanford University, focusing on regulatory DNA. This dataset contains 2.3 million regulatory DNA sequences, designed to assess the performance of DNA language models across zero-shot, probe-based, and fine-tuning settings. The dataset is constructed using curated cis-regulatory elements (cCREs) from the ENCODE Project and synthetic negative samples generated by preserving dinucleotide frequencies. Applications of DART-Eval cover sequence motif discovery, cell type-specific regulatory activity prediction, and counterfactual prediction of regulatory genetic variants, aiming to address the complexity of regulatory DNA sequences.

- 1DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA斯坦福大学 · 2024年



