遇见数据集

Art-historical Dataset with Structured Metadata and Iconographic Triplets from Wikidata

收藏
Zenodo2026-02-21 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains metadata for 36,043 artworks extracted from Wikidata. Using the Wikidata SPARQL endpoint, we retrieved all objects that are instances of either “visual artwork” (Wikidata item Q4502142) or “artwork series” (Q15709879). Each item in the dataset meets two conditions: It includes at least one two-dimensional digital image (P18), and It includes at least one Iconclass notation (P1257). For every object, we collected the following Wikidata properties: image (P18) title (P1476) creator (P170) inception (P571) instance of (P31) width (P2049) height (P2048) made from material (P186) movement (P135) genre (P136) depicts (P180) main subject (P921) depicts Iconclass notation (P1257) collection (P195) location (P276) country (P17) country of origin (P495) In addition to the Wikidata metadata, the dataset provides automatically generated concepts, tuples, and triplets derived from the Iconclass notations. These were produced using the Qwen3 large language model with the prompting template provided in the file prompt.txt. File structure The images are stored in a ZIP file structured into directories named by the first two characters of each image's hash_id. Within these directories, subfolders named after the next two characters of the hash_id contain the image files, which are named using their full hash_id with a .jpg extension. The annotation data is provided in a JSONL file, where each line encodes metadata for a single image.

提供机构:
Zenodo
创建时间:
2025-10-26
二维码
社区交流群
二维码
科研交流群
商业服务