遇见数据集

NancyT/clean_data2

收藏
Hugging Face2026-05-09 更新2026-05-31 收录
官方服务:

资源简介:

一个用于机器翻译或序列到序列任务的数据集,包含源语言和目标语言文本,以及token长度、比例等统计信息,训练集有158,382,105个样本。

A dataset for machine translation or sequence-to-sequence tasks, containing source and target language texts with token lengths, ratios, and other statistics, with 158,382,105 training examples.

提供机构:
NancyT
二维码
社区交流群
二维码
科研交流群
商业服务