遇见数据集

CJaFr-v3

收藏
arXiv2022-08-28 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

CJaFr-v3是一个免费提供的日法双语对齐语料库,包含1500万个对齐段落,由多个现有资源编译和过滤而成。该数据集涵盖了多种文本类型,如演讲、法律文本和百科全书式内容,旨在通过提供高质量的双语材料,促进日法语言对的翻译和分析研究。数据集的创建过程涉及对原始资源的筛选和处理,以确保内容的质量和适用性。CJaFr-v3的应用领域主要集中在机器翻译和语言分析,旨在解决日法语言对资源稀缺的问题。

CJaFr-v3 is a freely available Japanese-French bilingual aligned corpus containing 15 million aligned paragraphs, compiled and filtered from multiple existing resources. This corpus covers a wide range of text types, including speeches, legal documents, and encyclopedic content. It aims to facilitate translation and analytical research for the Japanese-French language pair by providing high-quality bilingual materials. The development of this corpus involves screening and processing of original resources to ensure the quality and applicability of the content. The application scenarios of CJaFr-v3 mainly focus on machine translation and language analysis, aiming to address the issue of scarce resources for the Japanese-French language pair.

创建时间:
2022-08-28
二维码
社区交流群
二维码
科研交流群
商业服务