遇见数据集

cskokgibbs/yeast-cv-split3

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个用于基因与转录因子相互作用预测的NLP数据集,包含基因标识符、转录因子名称、相互作用标签以及文本格式化输入和对应的token ID序列。特征包括:gene(基因)、TF(转录因子)、interaction(相互作用,整数值)、formatted_inputs(格式化输入文本)、input_ids(输入ID序列)、attention_mask(注意力掩码)和labels(标签)。数据集仅包含训练集,共有235,560个样本,总大小约1.03 GB,适用于自然语言处理任务,如序列分类或生成。

This dataset is an NLP dataset for predicting gene and transcription factor interactions, containing gene identifiers, transcription factor names, interaction labels, formatted text inputs, and corresponding token ID sequences. Features include: gene (gene), TF (transcription factor), interaction (interaction, integer value), formatted_inputs (formatted input text), input_ids (input ID sequence), attention_mask (attention mask), and labels (label). The dataset includes only a training set with 235,560 samples, total size approximately 1.03 GB, suitable for natural language processing tasks such as sequence classification or generation.

提供机构:
cskokgibbs
二维码
社区交流群
二维码
科研交流群
商业服务