fineweb_100BT-shuffled
收藏资源简介:
该数据集包含约1.6亿条训练样本,总规模达566GB,是一个多语言文本数据集。数据特征包括:原始文本内容(text)、唯一标识符(id)、数据来源信息(dump)、网页URL(url)、抓取日期(date)、文件路径(file_path)、语言标识(language)、语言置信度评分(language_score)、文本token计数(token_count)以及来源子数据集标识(dataset)。数据集适用于多语言文本处理、网络内容分析、语言识别等任务,其包含的质量评分和元数据可用于数据筛选和预处理。
This dataset contains approximately 160 million training samples with a total size of 566 GB, and it is a multilingual text dataset. Its data features include: raw text content (text), unique identifier (id), data source information (dump), webpage URL (url), crawl date (date), file path (file_path), language identifier (language), language confidence score (language_score), text token count (token_count), and source sub-dataset identifier (dataset). This dataset is applicable to tasks such as multilingual text processing, web content analysis, and language recognition. The included quality scores and metadata can be used for data filtering and preprocessing.



