Patent-CR
收藏资源简介:
Patent-CR是由剑桥大学创建的专利声明修订任务的首个数据集,包含22,606对初始申请和最终授权的专利声明。数据集内容涵盖了专利声明的修订过程,包括内容修正、术语一致性、语言精确性、简洁性和重新编号等五个主要修订类型。数据集的创建过程包括从Google Patents和欧洲专利局获取数据,并通过API提取和整理。该数据集主要应用于人工智能和自然语言处理领域,旨在解决专利声明修订中的复杂性和法律准确性问题。
Patent-CR is the first dataset for patent claim revision tasks, developed by the University of Cambridge. It contains 22,606 pairs of initial application and final granted patent claims. The dataset covers the entire revision process of patent claims, including five core revision types: content correction, terminology consistency, linguistic accuracy, conciseness, and renumbering. The dataset was built by collecting data from Google Patents and the European Patent Office, followed by API-based extraction and curation. It is primarily applied in the fields of artificial intelligence and natural language processing, with the goal of addressing the complexity and legal accuracy issues encountered during patent claim revision.




