相关数据集
Application of Chinese Pre-trained Language Models in Early Detection of Cognitive Impairment: A Comparative Study Based on Spoken Text
The dataset of \"Application of Chinese Pre-trained Language Models in Early Detection of Cognitive Impairment: A Comparative Study Based on Spoken Text\"
DataONE2025-12-07 更新220
CMLI-NLP/Mongolian-pretrain-dataset
--- license: cc-by-4.0 language: - mn size_categories: - 100K<n<1M --- # Mongolian Pretraining Dataset ## Dataset Information - **Language**: Mongolian (Traditional Mongolian script) - **Size**: ~1
Hugging Face2025-08-03 更新150
C4 (Colossal Clean Crawled Corpus)
C4 是 Common Crawl 的网络爬虫语料库的一个巨大的、干净的版本。它基于 Common Crawl 数据集:https://commoncrawl.org。它用于训练 T5 文本到文本的 Transformer 模型。可以从 allennlp 以预处理的形式下载数据集。
OpenXLab110
datajuicer/redpajama-arxiv-refined-by-data-juicer
RedPajama -- ArXiv数据集是经过Data-Juicer精炼的ArXiv数据集版本,通过移除一些低质量样本以提高数据集质量,通常用于预训练大型语言模型。该数据集包含1,655,259个样本,保留了原始数据集约95.99%的内容。精炼过程包括多种数据处理操作,如清理电子邮件和链接、修复Unicode、标点符号和空白标准化、过滤字母数字、平均行长度、字符重复、标记词、最大行长度、困惑度、
Hugging Face2023-10-23 更新150
Dudep/types-xlnet-base-cased
该数据集包含三个主要特征:labels、input_ids和attention_mask。labels特征是一个序列,包含27个不同的类别标签,每个标签代表一个特定的组合(如0-0-0、0-0-E3等)。input_ids和attention_mask特征都是int64类型的序列。数据集被分为训练集、验证集和测试集,分别包含120335、13371和14857个样本,对应的字节大小分别为37111
Hugging Face2024-07-05 更新140



