遇见数据集

Clustering and duplicates identification in Paintings on Wikidata

收藏
Zenodo2026-06-05 更新2026-06-12 收录
官方服务:

资源简介:

This dataset is the results of a study which explores how a relatively simple clustering method can then be used to identify duplicates and potential mistakes in the metadata of an already existing database. In this case, we studied the paintings on Wikidata, an open and collaborative knowledge base owned by the Wikimedia foundation. Clusters were generated on embeddings of the painting generated using DINOV3. Duplicates images were identified using Dense point matching. Further information on the method can be found on [LINK TO DATA PAPER IF PUBLISHED] images.tar, metadata.tar and embeddings.tar are archives containing the images for the paintings, the medadata from Wikidata in json format and the embeddings from DinoV3 respectively. dbscan_labels.csv describes the label produced by the DBSCAN clustering for each embedding. cluster_description.csv describes each unique cluster. duplicates.csv describes the results of hand-labeling of duplicates images (an thus of potential mistakes).

提供机构:
Zenodo
创建时间:
2026-06-05
二维码
社区交流群
二维码
科研交流群
商业服务