遇见数据集

textcleanlm/textclean-2B-raw-sample

收藏
Hugging Face2025-08-14 更新2025-09-13 收录
官方服务:

资源简介:

该数据集是一个文本数据集,包含了文本的ID、文本内容、元数据(如URL、来源域名、Warc相关信息等)、质量信号、FastText特征、eai_taxonomy分类信息等。数据集被划分为训练集,共包含100个文本示例。

This dataset is a collection of text data, including text ID, text content, metadata (such as URL, source domain, Warc-related information, etc.), quality signals, FastText features, eai_taxonomy classification information, etc. The dataset is divided into a training set, containing a total of 100 text examples.

提供机构:
textcleanlm
二维码
社区交流群
二维码
科研交流群
商业服务