遇见数据集

FRAME: A Dataset for Fine-grained Recognition of Art-historical Metadata and Entities

收藏
Zenodo2026-03-12 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains 200 manually annotated and curated art-historical image descriptions designed for Named Entity Recognition (NER) and Relation Extraction (RE). The descriptions were harvested from museum catalogs, auction listings, open-access platforms, and scholarly databases. Each record in the dataset includes the referenced artwork image, basic artwork metadata, and an art-historical text excerpt labeled with entity spans and typed relations. The dataset provides stand-off annotations in three layers: the metadata layer includes entity types that describe the artwork as an object (its material and production context); the content layer includes entity types that describe what the artwork represents (its figures and iconographic motifs); the co-reference layer includes links to repeated mentions of the same entity across the description. Across layers, entity mentions are labeled with 37 entity types and connected by typed relation links between mentions. Entity types are aligned with Wikidata. File structure The dataset is distributed as an INCEpTION export, with one directory per document. Document directory names follow the normalized schema {title}_{creator}.txt. Each document directory contains: CURATION_USER.ser (curation state) inception-document{id}.zip (annotation package) TypeSystem.xml (UIMA type system) CURATION_USER.xmi (curated UIMA XMI CAS) The referenced artwork images are stored in a ZIP file. In addition, the dataset includes a CSV metadata file with bibliographic information for each document, providing: document directory, title, creator, artwork creation date, source institution, and a URL to the original record.

提供机构:
Zenodo
创建时间:
2026-02-22
二维码
社区交流群
二维码
科研交流群
商业服务