遇见数据集

ardauzunoglu/c4_lowq_200m2b_subsample20m

收藏
Hugging Face2026-05-19 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含100,163个训练样本,每个样本具有文本内容、词元计数、文本质量评估指标(如低质量概率、得分和标签)以及来源URL。数据集可能用于文本质量分析、过滤或机器学习任务,总大小约为105 MB,下载大小约为62 MB。

This dataset contains 100,163 training examples, each with text content, token count, text quality assessment metrics (such as low-quality probability, score, and label), and source URL. It may be used for text quality analysis, filtering, or machine learning tasks, with a total size of approximately 105 MB and a download size of approximately 62 MB.

提供机构:
ardauzunoglu
二维码
社区交流群
二维码
科研交流群
商业服务