遇见数据集

ImageNet-GLT

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

对于ImageNet数据集,没有ground-truth属性标签可供我们监视类内分布并构建GLT基准。但是,在所有类型的数据中,普遍存在属性的长尾分布,即类内不平衡。因此,我们可以应用聚类算法 (例如项目中的K-Means) 在每个类中创建一组聚类作为 “借口属性”。换句话说,每个集群代表这个类的元属性布局,例如,一个用于猫类别的集群可以是ginger cat In house,另一个用于猫类别的集群是black cat on street等。 请注意,我们这里的聚类属性纯粹是基于类内特征分布构建的,因此它代表了导致类内方差的各种因素,包括对象级属性 (如纹理) 或图像级属性 (如上下文)。

For the ImageNet dataset, there are no ground-truth attribute labels available for us to monitor intra-class distribution and construct the GLT benchmark. However, the long-tail distribution of attributes universally exists across all types of data, namely intra-class imbalance. Therefore, we can apply clustering algorithms (e.g., K-Means in this project) to create a set of clusters within each class as "excuse attributes". In other words, each cluster represents the meta-attribute layout of the corresponding class. For example, one cluster for the cat category could be "ginger cat in house", while another could be "black cat on street", and so on. Please note that the clustering attributes used here are purely constructed based on intra-class feature distribution, thus they represent various factors leading to intra-class variance, including object-level attributes (e.g., texture) or image-level attributes (e.g., context).

提供机构:
OpenDataLab
创建时间:
2022-11-02
搜集汇总
数据集介绍
ImageNet-GLT 数据集图片
背景与挑战
背景概述
ImageNet-GLT是基于ImageNet数据集构建的广义长尾基准,通过聚类算法在各类别内创建借口属性来模拟类内不平衡分布。该数据集由南洋理工大学、浙江大学和阿里巴巴达摩学院于2022年联合发布。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务