遇见数据集

NancyT/clean_data

收藏
Hugging Face2026-05-09 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含文本、长度、损失、源语言和目标语言的token长度、总token长度、token比例、词比例、目标语言、源语言和模式分数等特征,用于机器翻译或文本生成任务。训练集包含约9479万个样本,总大小约23.6GB。

This dataset includes features such as text, length, loss, source token length, target token length, total token length, token ratio, word ratio, target language, source language, and pattern score, suitable for machine translation or text generation tasks. The training set contains approximately 94.79 million samples with a total size of about 23.6 GB.

提供机构:
NancyT
二维码
社区交流群
二维码
科研交流群
商业服务