SIRBOT - Semantic Image Retrieval Based on Object Translation
收藏资源简介:
There is an urgent need to develop an efficient and effective semantic based image retrieval (SBIR) tool that has similar functionality and flexibility like textual document retrieval tools. In this thesis, we propose the framework for an SBIR system. The main idea of the proposed system is to translate images into textual documents so that images can be retrieved the same way as the textual documents are retrieved. The system breaks images down into regions and represents these regions with effective colour and texture features. For each type of feature, we generate a visual dictionary which is used to discretize continuous valued features into a finite set of representative features. In the next, semantic concepts are learnt for the regions using a decision tree (DT) based machine learning technique. Finally, the images are indexed using an inverted file and retrieved by keywords. Several innovations have been achieved during this research. First, a system of semantic image retrieval based on object translation (SIRBOT) is proposed. The system is the first of this kind to conceptualize a complete SBIR system. Second, a new directionality feature is proposed based on the geometric property of the edge histogram of an image. Third, a shape transform method is proposed to apply rotation invariant curvelet transform on irregular colour regions. Fourth, an adaptive vector quantization (AVQ) algorithm is proposed to generate a visual dictionary from a set of training region features by quantizing the feature space. The proposed AVQ is adaptive to both variable dimension feature vectors and data size. Fifth, a region based inverted file system is proposed to index the translated images. The characteristics of image documents are exploited to adapt the text retrieval technique in semantic image retrieval. Experimental results show that the proposed SIRBOT system outperforms both the low level retrieval and the widely used Bayesian models.



