Custom dataset
收藏资源简介:
本研究介绍了从网络数据中无监督生成多标签数据集的系统。该数据集名为Custom dataset,由Vilynx Inc.创建,包含250,000张图像,涉及96种不同场景和地点的类别。数据集的创建过程包括使用图像搜索引擎收集样本,并通过聚类和锚点选择进行噪声降低。此数据集主要用于图像分类任务,旨在解决现有数据集大小和多样性不足的问题,特别是在工业应用中。
This study introduces a system for unsupervised generation of multi-label datasets from web data. Named Custom dataset, the dataset was developed by Vilynx Inc. and consists of 250,000 images spanning 96 categories of distinct scenes and locations. The dataset creation workflow includes collecting samples through image search engines, followed by noise reduction via clustering and anchor selection. This dataset is mainly applied to image classification tasks, aiming to address the shortcomings of existing datasets in terms of scale and diversity, especially in industrial applications.



