遇见数据集

cskokgibbs/yeast-cv-split2

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含基因与转录因子之间的相互作用数据,用于生物信息学或基因调控分析任务。特征包括基因名称(gene)、转录因子名称(TF)、相互作用标签(interaction)、格式化输入文本(formatted_inputs)、输入标识符列表(input_ids)、注意力掩码(attention_mask)和标签(labels)。数据集适用于机器学习或自然语言处理模型训练,可能用于预测基因与转录因子之间的相互作用类型或强度。数据已分割为训练集,包含184,080个样本,总大小约为777.8 MB。

This dataset contains interaction data between genes and transcription factors, intended for bioinformatics or gene regulation analysis tasks. Features include gene name (gene), transcription factor name (TF), interaction label (interaction), formatted input text (formatted_inputs), input identifier list (input_ids), attention mask (attention_mask), and labels (labels). The dataset is suitable for training machine learning or natural language processing models, potentially for predicting the type or strength of interactions between genes and transcription factors. The data is split into a training set with 184,080 samples and a total size of approximately 777.8 MB.

提供机构:
cskokgibbs
二维码
社区交流群
二维码
科研交流群
商业服务