HuggingFaceBio/malinois-mpra-regression
收藏资源简介:
该数据集对Gosai等人2024年发表的Nature论文中的补充MPRA表格进行了预处理,用于Malinois基准中的监督DNA到活性回归任务。每一行包含一个DNA序列和三个细胞类型特异性活性目标:K562_log2FC、HepG2_log2FC和SKNSH_log2FC。数据集规模在10万到100万行之间,适用于表格回归任务,涉及生物学、基因组学、DNA和MPRA领域。数据分割遵循Carbon微调实验和公开的Malinois设置,包括训练集、验证集和测试集,总行数经过过滤后为798,064行。列包括ID、分割标签、染色体信息、元数据、DNA序列、反向互补序列、连接序列、原始回归目标、标准误差、标准化目标以及过滤标志等。使用示例展示了如何加载数据集并提取序列和标签。
This dataset preprocesses the Gosai et al. 2024 supplementary MPRA table used by the Malinois benchmark for supervised DNA-to-activity regression. Each row contains a DNA sequence and three cell-type-specific activity targets: K562_log2FC, HepG2_log2FC, and SKNSH_log2FC. The dataset size is between 100K and 1M rows, categorized under tabular-regression tasks with tags in biology, genomics, dna, and mpra. Splits follow the Carbon fine-tuning experiments and the public Malinois setup, including train, validation, and test sets, with a total of 798,064 rows after filtering. Columns include ID, split, chromosome, metadata, DNA sequence, reverse complement, concatenated sequence, raw regression targets, standard errors, standardized targets, and filtering flags. Usage examples demonstrate loading the dataset and extracting sequences and labels.




