TaoTern/TaoData
收藏资源简介:
TaoData是一个为Taotern模型系列准备的大规模英文预训练数据集。它是HuggingFaceFW/fineweb-edu数据集的经过处理的子集,使用TaoData框架通过taodata.yaml配置进行过滤和格式化。该版本包含约670万个JSONL格式的样本,旨在作为Taotern模型使用的主要预训练语料库之一。数据集支持因果语言建模、继续预训练和基础模型的数据混合构建等任务。数据为英文,每个样本是一个JSON对象,包含文本、ID、来源URL、语言标签、质量分数等字段。数据集源自Common Crawl,经过清理和预处理,适用于大规模语言模型预训练和研究,但可能存在网络文本的局限性,如事实错误、偏见或噪声。
TaoData is a large-scale English pretraining dataset prepared for the Taotern model family. It is a processed subset of HuggingFaceFW/fineweb-edu, filtered and formatted with the TaoData framework using the taodata.yaml configuration. This release contains about 6.7 million samples in JSONL format and is intended as one of the main pretraining corpora used by Taotern models. It supports tasks such as causal language modeling, continued pretraining, and data mixture construction for foundation models. The data is in English, with each sample as a JSON object containing fields like text, ID, source URL, language tags, and quality scores. Derived from Common Crawl, it is cleaned and preprocessed for large-scale language model pretraining and research, but may have limitations typical of web-derived text, such as factual errors, bias, or noise.



