登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
RTM-20M Research Corpus v3
RTM-20M Research Corpus v3
收藏
kaggle
2026-07-26 更新
2026-08-01 收录
语言模型预训练
记忆增强训练语料
数据链接:
https://www.kaggle.com/datasets/nexuseval/rtm20m-research-corpus-v3
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
400M pre-training tokens with RTM memory and post-training data
应用场景:
创建时间:
2026-07-26
相关数据集
datajuicer/the-pile-uspto-refined-by-data-juicer
语言模型预训练
数据精炼
--- license: apache-2.0 task_categories: - text-generation language: - en tags: - data-juicer - pretraining size_categories: - 1M<n<10M --- # The Pile -- USPTO (refined by Data-Juicer) A refined ver
Hugging Face
2023-10-23 更新
39
0
pietrolesci/pile-deduped-pythia-preshuffled
语言模型预训练
文本语料库
这个数据集包含了完全准备好的数据,这些数据已经被标记化并预洗牌,用于训练Pythia(去重)模型。该数据集与EleutherAI组织下的EleutherAI/pile-deduped-pythia-preshuffled数据集相同,但以更易于管理的格式呈现。数据集分为143个块(parquet文件),每个块包含1024000个序列(行),对应1000个批次,每个批次由1024个序列组成。数据集包含
Hugging Face
2025-03-25 更新
10
0
parameterlab/scaling_mia_the_pile_00_NIH_ExPorter
学术文本挖掘
语言模型预训练
--- dataset_info: features: - name: text dtype: string - name: meta struct: - name: pile_set_name dtype: string splits: - name: train num_bytes: 129283623 num_examp
Hugging Face
2024-09-24 更新
5
0
TiWu-Lab/C4-zh
中文语料清洗
语言模型预训练
--- license: odc-by --- Chinese text cleaned from [C4](https://huggingface.co/datasets/allenai/c4) with the following steps: - documents containing non-Chinese, non-English text are removed - docume
Hugging Face
2025-03-29 更新
20
0
NeelNanda/c4-tokenized-2b
语言模型预训练
文本分词数据
--- dataset_info: features: - name: tokens sequence: int64 splits: - name: train num_bytes: 11145289620 num_examples: 1359845 download_size: 2530851147 dataset_size: 1114528962
Hugging Face
2022-11-14 更新
8
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广