遇见数据集

KantaHayashiAI/ClimbLab-Ja

收藏
Hugging Face2026-05-13 更新2026-05-03 收录
官方服务:

资源简介:

ClimbLab-Ja是一个经过过滤的3000亿令牌日语语料库,包含20个集群。它是基于LLM-jp Corpus v4的日语适配版本,通过语义重组和过滤形成高质量语料库。具体来说,首先根据主题信息将数据分组为1000个组,然后为每个组和文档分配六个评分(0-5分):质量、广告、信息价值、教育价值、文化价值和创意价值。基于这些评分,移除了低质量的集群和文档。该数据集仅用于研究和开发,适用于预训练语言模型,格式为parquet文本。

ClimbLab-Ja is a filtered 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group and document: quality, advertisement, informational value, educational value, cultural value, and creative value. Low-quality clusters and documents were removed based on these scores. This dataset is for research and development only, intended for pre-training language models, and is in parquet format.

提供机构:
KantaHayashiAI
二维码
社区交流群
二维码
科研交流群
商业服务