JaParaPat
收藏资源简介:
JaParaPat是由NTT Corporation创建的一个大规模的日英平行专利申请语料库,包含超过3000万日英句对,数据来自2000年至2021年间在日本和美国发布的专利申请。数据集内容涵盖了专利申请的标题、摘要、描述和权利要求等部分,通过基于翻译的句子对齐方法进行提取,数据集的创建过程包括从日本专利局(JPO)和美国专利商标局(USPTO)获取未审查的专利申请,以及从欧洲专利局(EPO)的DOCDB数据库获取专利家族信息。JaParaPat旨在解决专利翻译中的质量问题,并用于研究和开发机器翻译技术。
JaParaPat is a large-scale Japanese-English parallel patent application corpus developed by NTT Corporation. It contains over 30 million Japanese-English sentence pairs, with data sourced from patent applications published in Japan and the United States between 2000 and 2021. The corpus covers core sections of patent applications including titles, abstracts, detailed descriptions and claims. It was extracted via translation-based sentence alignment methods. The construction of JaParaPat involves acquiring unexamined patent applications from the Japan Patent Office (JPO) and the United States Patent and Trademark Office (USPTO), as well as patent family information from the DOCDB database of the European Patent Office (EPO). JaParaPat aims to address quality issues in patent translation, and is intended for research and development of machine translation technologies.



