Multi-view multi-label canonical correlation analysis for cross-modal multimedia retrieval
摘要
We address the problem of cross-modal retrieval in presence of multi-view and multi-label data. For this, we present Multi-view Multi-label Canonical Correlation Analysis (or MVMLCCA), which is a generalization of Canonical Correlation Analysis (CCA) for multi-view data. While CCA relies on explicit pairings/associations of samples between two views (or modalities), MVMLCCA uses high-level semantic information (such as multi-label annotations) to establish correspondence across multiple (two or more) views without the need of an explicit pairing/association of samples across multiple views. We also present Fast MVMLCCA (FMVMLCCA), a computationally efficient version of MVMLCCA which is experimentally shown to be several times faster than MVMLCCA. Extensive experimental analyses on three popular multi-modal datasets (IAPRTC-12, MS-COCO and MIR-Flickr) demonstrate that the proposed approach offers much more flexibility than the related approaches without compromising on scalability and cross-modal retrieval performance.