遇见数据集

MK0727/noise-line-label-jp

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

noise-line-label-jp是一个小型日语数据集,源自`HuggingFaceFW/fineweb-2`。它包含源文本样本和行级标注,这些标注用于识别在高质量预训练语料库中建议保留的行。数据集结构包括:id(来自FineWeb2的源文档ID)、text(原始日语文本样本)和lines_to_keep(建议保留的1索引行号)。该数据集旨在用于日语语料库清理、行级噪声检测以及语言模型训练数据的预处理工作流实验。文本样本取自FineWeb2的日语分割部分,保留标签由LLM自动生成,在生产使用前应进行审查。

noise-line-label-jp is a small Japanese dataset derived from `HuggingFaceFW/fineweb-2`. It contains source text samples and line-level annotations that identify lines recommended for keeping in a higher-quality pre-training corpus. The dataset structure includes: id (source document ID from FineWeb2), text (original Japanese text sample), and lines_to_keep (1-indexed line numbers recommended for keeping). It is intended for experimenting with Japanese corpus cleaning, line-level noise detection, and preprocessing workflows for language model training data. The text samples are taken from the Japanese split of FineWeb2, and the keep labels were generated automatically by LLM and should be reviewed before production use.

提供机构:
MK0727
二维码
社区交流群
二维码
科研交流群
商业服务