登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
Performance comparison of DNABERT with and without pre-training.
Performance comparison of DNABERT with and without pre-training.
收藏
Figshare
2025-11-12 更新
2026-04-28 收录
DNA序列分析
预训练模型
数据链接:
https://figshare.com/articles/dataset/Performance_comparison_of_DNABERT_with_and_without_pre-training_/30603036
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
Performance comparison of DNABERT with and without pre-training.
应用场景:
创建时间:
2025-11-12
相关数据集
datajuicer/llava-pretrain-refined-by-data-juicer
多模态学习
预训练模型
LLaVA pretrain -- LCS-558k数据集是LLaVA预训练数据集的一个精炼版本,通过Data-Juicer工具去除了原始数据集中的一些低质量样本,以提高数据集的质量。该数据集主要用于多模态大语言模型的预训练。数据集的样本数量为500,380个,保留了原始数据集的约89.65%。精炼过程包括修复Unicode错误、标点符号规范化、过滤不符合特定条件的文本和图像样本等步骤。
Hugging Face
2024-03-07 更新
129
0
ih138/my_awesome_eli5_mlm_model
自然语言处理
预训练模型
该数据集包含两个主要特征:input_ids和attention_mask,分别表示为int32和int8类型的序列。数据集分为训练集和测试集,训练集包含4000个样本,测试集包含1000个样本。数据集的下载大小为3587573字节,总大小为8464880字节。数据文件路径分别为data/train-*和data/test-*。
Hugging Face
2024-07-10 更新
12
0
seungone/ablation1_math_gpt4o_mini
数学文本处理
预训练模型
该数据集包含两个主要特征:instruction(指令)和response(响应),均为字符串类型。数据集包含一个训练集(train),其中包含5561个样本,总大小为7702223字节。数据集的下载大小为3515007字节。数据集的配置信息显示,训练集数据文件位于路径data/train-*。
Hugging Face
2024-11-25 更新
8
0
Google Word2Vec
自然语言处理
预训练模型
Pre-trained word to vector model
kaggle
2021-09-17 更新
17
0
SH1045689.10FU
真菌分类学
DNA序列分析
UNITE provides a unified way for delimiting, identifying, communicating, and working with DNA-based Species Hypotheses (SH). All fungal ITS sequences in the international nucleotide sequence databases
DataCite Commons
2024-09-05 更新
9
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广