登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
checkpoint of alchemBERT for matbench-discovery
checkpoint of alchemBERT for matbench-discovery
收藏
Figshare
2025-03-29 更新
2026-04-08 收录
化学材料发现
预训练语言模型
数据链接:
https://figshare.com/articles/dataset/checkpoint_of_alchemBERT_for_matbench-discovery/28690583/1
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
checkpoint of alchemBERT for matbench-discovery
应用场景:
提供机构:
Wang, Yuhang
创建时间:
2025-03-29
相关数据集
C4 (Colossal Clean Crawled Corpus)
文本语料库
预训练语言模型
C4 是 Common Crawl 的网络爬虫语料库的一个巨大的、干净的版本。它基于 Common Crawl 数据集:https://commoncrawl.org。它用于训练 T5 文本到文本的 Transformer 模型。可以从 allennlp 以预处理的形式下载数据集。
OpenXLab
20
0
Do Pre-trained Language Models Indeed Understand Software Engineering Tasks?
预训练语言模型
软件工程任务
Datasets and code in "Do Pre-trained Language Models Indeed Understand Software Engineering Tasks?"
NIAID Data Ecosystem
9
0
mhr2004/plm-train-nsp-1000000-tokenized
自然语言处理
预训练语言模型
--- dataset_info: features: - name: text dtype: string - name: label dtype: int64 - name: input_ids sequence: int32 - name: attention_mask sequence: int8 splits: - name:
Hugging Face
2024-04-30 更新
8
0
CMLI-NLP/Mongolian-pretrain-dataset
蒙古语处理
预训练语言模型
--- license: cc-by-4.0 language: - mn size_categories: - 100K<n<1M --- # Mongolian Pretraining Dataset ## Dataset Information - **Language**: Mongolian (Traditional Mongolian script) - **Size**: ~1
Hugging Face
2025-08-03 更新
15
0
BiodivBERT: Pre-training Corpora DOIs
生物多样性文本挖掘
预训练语言模型
This repository contains the DOIs we used to construct the pre-training corpora for BiodivBERT model. BiodivBertAbs uses the abstracts DOIs while BiodivBERTAbs+Full uses both of them.
NIAID Data Ecosystem
9
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广