Detailed information of each dataset.
收藏资源简介:
Multi-visual pattern mining plays an important role in image classification, retrieval, and other fields. A multi visual pattern mining algorithm based on variational inference Gaussian mixture model and pattern activation response graph is introduced to address the issues of insufficient frequency and discriminability faced by traditional algorithms. The innovation of this algorithm lies in combining variational inference Gaussian mixture model with pattern activation response graph. The former solves the limitation of manually presetting the number of modes in traditional methods by determining the optimal number of modes to ensure frequency. The latter improves discriminability by capturing key areas of the image, solving the problem of traditional algorithms being difficult to balance the two and distinguish multiple patterns within the same category. The results showed that in quantitative analysis, the algorithm had a high frequency of 92.81% when the similarity threshold was 0.866 on the Canadian Institute for Advanced Research-10 dataset. On the Travel dataset, the classification accuracy and F1 value were as high as 95.36% and 94.17%, respectively, which were significantly higher than other algorithms. The proposed multi-visual pattern mining algorithm has high frequency and discriminability, which can provide a more comprehensive visual representation and help better mine images of the same category but different visual patterns. This algorithm provides technical support for image classification and retrieval.
多视觉模式挖掘在图像分类、检索等领域发挥着关键作用。针对传统算法存在的频率性与可区分性不足的问题,本文提出一种基于变分推理高斯混合模型(variational inference Gaussian mixture model)与模式激活响应图(pattern activation response graph)的多视觉模式挖掘算法。该算法的创新点在于将上述两种方法相结合:前者通过确定最优模式数量,规避了传统方法中手动预设模式数的局限,以此保障频率性;后者则通过捕获图像关键区域提升可区分性,解决了传统算法难以兼顾二者且无法区分同一类别内多种视觉模式的难题。实验结果表明,在定量分析中,于加拿大高级研究学院10(Canadian Institute for Advanced Research-10)数据集上,当相似度阈值为0.866时,该算法的频率性可达92.81%;在Travel数据集上,其分类准确率与F1值分别高达95.36%与94.17%,显著优于其他对比算法。所提多视觉模式挖掘算法兼具优异的频率性与可区分性,能够提供更全面的视觉表征,助力更好地挖掘同一类别内具备不同视觉模式的图像,可为图像分类与检索任务提供技术支撑。




