联合特征精炼与多模态分类器学习的地标图像检索

Landmark Image Retrieval by Jointing Feature Refinement and Multimodal Classifier Learning

IEEE Transactions on Cybernetics · 2017
被引 16
ABS 3

中文导读

提出一种联合特征精炼和多模态分类器学习的方法,利用社交图像的地理相关性进行地标检索,通过低秩矩阵恢复精炼视觉特征,结合组稀疏多模态分类,在真实数据集上优于现有方法。

Abstract

Landmark retrieval is to return a set of images with their landmarks similar to those of the query images. Existing studies on landmark retrieval focus on exploiting the geometries of landmarks for visual similarity matches. However, the visual content of social images is of large diversity in many landmarks, and also some images share common patterns over different landmarks. On the other side, it has been observed that social images usually contain multimodal contents, i.e., visual content and text tags, and each landmark has the unique characteristic of both visual content and text content. Therefore, the approaches based on similarity matching may not be effective in this environment. In this paper, we investigate whether the geographical correlation among the visual content and the text content could be exploited for landmark retrieval. In particular, we propose an effective multimodal landmark classification paradigm to leverage the multimodal contents of social image for landmark retrieval, which integrates feature refinement and landmark classifier with multimodal contents by a joint model. The geo-tagged images are automatically labeled for classifier learning. Visual features are refined based on low rank matrix recovery, and multimodal classification combined with group sparse is learned from the automatically labeled images. Finally, candidate images are ranked by combining classification result and semantic consistence measuring between the visual content and text content. Experiments on real-world datasets demonstrate the superiority of the proposed approach as compared to existing methods.

地标检索多模态学习图像分类特征精炼社交图像