Clustering and duplicates identification in Paintings on Wikidata
收藏资源简介:
This dataset is the results of a study which explores how a relatively simple clustering method can then be used to identify duplicates and potential mistakes in the metadata of an already existing database. In this case, we studied the paintings on Wikidata, an open and collaborative knowledge base owned by the Wikimedia foundation. Clusters were generated on embeddings of the painting generated using DINOV3. Duplicates images were identified using Dense point matching. Further information on the method can be found on [LINK TO DATA PAPER IF PUBLISHED] images.tar, metadata.tar and embeddings.tar are archives containing the images for the paintings, the medadata from Wikidata in json format and the embeddings from DinoV3 respectively. dbscan_labels.csv describes the label produced by the DBSCAN clustering for each embedding. cluster_description.csv describes each unique cluster. duplicates.csv describes the results of hand-labeling of duplicates images (an thus of potential mistakes).



