遇见数据集

PAIR DATASET AND ANNOTATION

收藏
Zenodo2026-02-09 更新2026-05-29 收录
官方服务:

资源简介:

PAIR: Panofskian Art Interpretation and Retrieval Framework - Datasets and Annotations This repository contains the datasets and pre-computed annotations for the paper "PAIR: A Panofsky-Inspired Multi-Agent Framework for Art Image Retrieval". Repository Contents File Size Description annotation-data.zip - Pre-computed multi-agent annotations for all artworks semart-images-1.zip - SemArt dataset images (Part 1) semart-images-2.zip - SemArt dataset images (Part 2) semart-files.zip - SemArt dataset metadata and query files 1. Annotation Data (annotation-data.zip) Pre-computed multi-agent annotations for all artworks in the evaluation datasets. Contents Structured annotations following Panofsky's three-level iconographic framework Global annotations (form, subject, symbolism, cultural context, technical analysis) Regional annotations with spatial coordinates Vector embeddings (BGE-large-en-v1.5, 1024 dimensions) Dataset Scales Annotations are provided for three evaluation scales: Dataset Images Purpose data-100 100 Rapid validation data-1069 1,069 Primary benchmark (complete test set) data-5000 5,000 Scalability testing 2. SemArt Dataset The SemArt benchmark is split into three files due to size constraints. Images semart-images-1.zip: Artwork images (Part 1) semart-images-2.zip: Artwork images (Part 2) Metadata (semart-files.zip) Image metadata (title, artist, date, type) Official train/test splits 100 evaluation queries with ground-truth annotations License SemArt is released under CC BY-NC 4.0 (Creative Commons Attribution-NonCommercial). See the original dataset page for details. All experimental datasets are derived from the publicly available SemArt benchmark. The 100 evaluation queries are drawn from SemArt's official test partition, testing iconographic understanding, symbolic interpretation, and cultural context comprehension. Image Sources: Artwork images in SemArt originate from museum digital collections including the Web Gallery of Art and Wikimedia Commons. Many source images are in the public domain due to artwork age (pre-1900). 3. Data Usage and Copyright Annotation Outputs PAIR generates original textual annotations through LLM-based analysis. These annotations constitute derivative scholarly analysis rather than reproduction of copyrighted content. Annotations describe visual elements, interpret symbolic meanings, and provide cultural context—all original analytical text generated by our system. No Image Reproduction Our system does not reproduce, modify, or redistribute copyrighted artwork images. During retrieval, the system returns references to original images (URLs) rather than copies. Users access images through their original sources. Model Outputs LLM-generated content (annotations and verification reasoning) represents original analytical work. These outputs do not reproduce protected expression from training data but synthesize domain knowledge into new scholarly descriptions.

提供机构:
Zenodo
创建时间:
2026-01-11
二维码
社区交流群
二维码
科研交流群
商业服务