遇见数据集

Statistical Metric Learning Benchmark (SMLB)

收藏
arXiv2023-12-02 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

Statistical Metric Learning Benchmark (SMLB) 是一个基于ImageNet-21K和WordNet构建的大规模数据集,旨在评估自监督学习模型的判别能力和泛化性。该数据集包含超过1400万张图像,分为20498个类别和16632个分类节点,覆盖了广泛的类别多样性和粒度。SMLB通过引入新的评估指标——‘overlap’和‘aSTD’,来衡量特征空间中不同类别间的可分离性和相似性分布的一致性。这些指标有助于更深入地理解自监督学习模型在处理复杂和不确定的现实世界问题时的表现,揭示了监督学习在数据集偏差和域迁移方面的局限性,为未来模型改进提供了方向。

Statistical Metric Learning Benchmark (SMLB) is a large-scale dataset built upon ImageNet-21K and WordNet, designed to evaluate the discriminative ability and generalization performance of self-supervised learning models. This dataset contains over 14 million images, categorized into 20,498 classes and 16,632 classification nodes, covering a wide spectrum of category diversity and granularity. SMLB introduces two novel evaluation metrics, 'overlap' and 'aSTD', to measure the separability between distinct categories in the feature space and the consistency of similarity distributions. These metrics facilitate a deeper comprehension of the performance of self-supervised learning models when addressing complex and uncertain real-world scenarios, reveal the limitations of supervised learning regarding dataset bias and domain adaptation, and offer actionable guidance for future model optimization.

提供机构:
萨里大学
创建时间:
2023-12-02
二维码
社区交流群
二维码
科研交流群
商业服务