遇见数据集

agentlans/en-document-classification

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

该数据集提供了来自allenai/c4(英文配置)的前100万行精选子集,并丰富了多视角主题标注。它专为研究文档分类、领域适应和大规模网络爬取语料库中的标签噪声的研究者设计。数据集整合了分类模型的预测结果,以提供每个文档内容的整体视图。为确保实用性,数据在这些变量上进行了分层,并分为80/10/10%(训练/验证/测试)集,根据源分类器组织成特定的配置。

This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. The dataset integrates predictions from classification models to provide a holistic view of each documents content. To ensure utility, the data was stratified across these variables and split into 80/10/10% (Train/Validation/Test) sets, organized into specific configurations based on the source classifier.

提供机构:
agentlans
二维码
社区交流群
二维码
科研交流群
商业服务