面向视觉识别与跨模态检索的广义多视角嵌入

Generalized Multi-View Embedding for Visual Recognition and Cross-Modal Retrieval

IEEE Transactions on Cybernetics · 2017
被引 108
ABS 3

中文导读

提出一个基于瑞利商的统一子空间学习框架,可扩展至多视角、监督学习和非线性嵌入,在视觉识别和跨模态图像检索中取得优于现有方法的结果。

Abstract

In this paper, the problem of multi-view embedding from different visual cues and modalities is considered. We propose a unified solution for subspace learning methods using the Rayleigh quotient, which is extensible for multiple views, supervised learning, and nonlinear embeddings. Numerous methods including canonical correlation analysis, partial least square regression, and linear discriminant analysis are studied using specific intrinsic and penalty graphs within the same framework. Nonlinear extensions based on kernels and (deep) neural networks are derived, achieving better performance than the linear ones. Moreover, a novel multi-view modular discriminant analysis is proposed by taking the view difference into consideration. We demonstrate the effectiveness of the proposed multi-view embedding methods on visual object recognition and cross-modal image retrieval, and obtain superior results in both applications compared to related methods.

计算机视觉模式识别机器学习跨模态检索子空间学习