新构建的多模态实体匹配语料库
收藏资源简介:
新构建的多模态实体匹配语料库,由中国信息处理实验室创建,包含超过120,000条记录,涉及10,000种产品。每条记录包含高质量的图像属性,用于研究多模态信息在实体匹配中的应用。数据集旨在通过移除限制性实体、平衡标签和单模态记录的假设,重新构建实体匹配基准,以更真实地反映实体匹配在开放环境中的挑战。该数据集适用于评估模型在处理未知实体集群、不平衡标签和多模态记录的能力,特别关注于解决实体匹配在实际应用中的性能问题。
A newly constructed multimodal entity matching corpus, developed by the China Information Processing Laboratory, contains over 120,000 records spanning 10,000 distinct products. Each record includes high-quality image attributes, supporting research on the application of multimodal information in entity matching tasks. This corpus aims to reconstruct entity matching benchmarks by removing restrictive entities, balancing label distributions and relaxing the assumption of unimodal records, thereby more realistically reflecting the challenges of entity matching in open environments. This dataset is suitable for evaluating model capabilities in handling unknown entity clusters, imbalanced labels and multimodal records, with a particular focus on addressing performance issues of entity matching in real-world applications.

- 1Bridging the Gap between Reality and Ideality of Entity Matching: A Revisiting and Benchmark Re-Construction中国信息处理实验室 · 2022年



