遇见数据集

Yokii2/kosuzu-jaen

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

Kosuzu JA-EN是一个大规模日语→英语平行翻译数据集,通过OPUS从两个开放资源(CCMatrix和Wikimedia)构建,并使用大型语言模型(LLMs)进行处理。数据集以东方Project中的Kosuzu Motoori命名。数据来源包括CCMatrix(占数据集的99.73%,主要使用Mistral Small 4模型处理)和Wikimedia(占0.27%,主要使用DeepSeek V3.2模型处理),部分数据保留原始对齐翻译(标记为passthrough)。数据集模式包括输入(原始日语文本)、输出(英语翻译)、来源(原始文件)和模型(使用的模型或passthrough)。数据规模在1亿到10亿条之间,属于合成类型,适用于翻译任务。

A large-scale Japanese → English parallel translation dataset, built from two open sources via OPUS and processed with LLMs. Named after Kosuzu Motoori from Touhou Project. Data sources include CCMatrix (99.73% of the dataset, primarily processed with Mistral Small 4) and Wikimedia (0.27%, primarily processed with DeepSeek V3.2), with some data kept as original aligned translations (marked as passthrough). The schema consists of input (original Japanese text), output (English translation), origin (source file), and model (model used or passthrough). The dataset size is between 100 million and 1 billion entries, is synthetic, and is intended for translation tasks.

提供机构:
Yokii2
二维码
社区交流群
二维码
科研交流群
商业服务