遇见数据集

CLEF-IP 2012

收藏
DataCite Commons2024-08-27 更新2024-07-13 收录
官方服务:

资源简介:

CLEF-IP: Cross-Language Evaluation Forum - Intellectual PropertyThe CLEF-IP track ran from 2009 to 2013 and aimed to investigate IR techniques for patent retrieval.The track utilizes a collection of more than 1.3M patent documents (~2.6 million files) derived from EPO (European Patent Office) sources and EuroPCT Applications (more than 400K documents) published by WIPO (World Intelectual Property Organization). The collection contains documents in English, French and German with at least 150,000 documents in each language, all published before 2001.There were three tasks in 2012: The first one was to find patent documents that are candidates to constitute prior art for a given claim taken from a patent document. The second task, flowchart recognition, asked participants to extract the information in these images and return it in a predefined textual format. The third task, chemical structure regonition, participants had to identify the location of the chemical structures depicted on images of patent pages and, for each of them, return the corresponding structure in a MOL file (a chemical structure file format).FilesDocument CollectionThe first one is a set of XML files representing a total of over 1.3 million patent documents.NOTE: the document collection is the same as the one published for CLEF-IP 2011, excluding images.Topics and AnswersBoth the training and the test topic sets contain also the relevance assessments for the topics.

CLEF-IP:跨语言评测论坛-知识产权(Cross-Language Evaluation Forum - Intellectual Property) CLEF-IP赛道于2009年至2013年间举办,旨在探索面向专利检索的信息检索技术。该赛道采用的数据集源自欧洲专利局(European Patent Office,EPO)公开的超130万件专利文档(约260万份文件),以及世界知识产权组织(World Intellectual Property Organization,WIPO)发布的EuroPCT申请文档(超40万件)。该数据集涵盖英语、法语、德语三种语言的文档,每种语言至少包含15万份,且所有文档均于2001年之前公开。 2012年该赛道设有三项任务:第一项任务为针对从专利文档中提取的给定权利要求,检索可作为其现有技术候选的专利文档;第二项任务为流程图识别,要求参与者从专利图像中提取信息并以预定义的文本格式返回结果;第三项任务为化学结构识别,参与者需定位专利页面图像中的化学结构,并为每个结构返回对应的MOL文件(一种化学结构文件格式)。 文件集:文档集合 第一部分为一组XML格式文件,总计涵盖超130万件专利文档。注:该文档集与CLEF-IP 2011发布的文档集一致,仅移除了图像文件。 主题与标注集 训练与测试主题集均包含对应主题的相关性标注结果。

提供机构:
TU Wien
创建时间:
2021-11-30
搜集汇总
数据集介绍
CLEF-IP 2012 数据集图片
背景与挑战
背景概述
CLEF-IP 2012是一个用于专利检索研究的测试集合,包含超过130万份多语言(英语、法语、德语)专利文档,所有文档均发表于2001年之前。该数据集支持三个主要任务:专利现有技术检索、流程图识别和化学结构识别,旨在评估信息检索技术在知识产权领域的应用。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务