Introduction
摘要
In this book, our emphasis is on multimodal information retrieval, specifically concentrating on text and image data. The traditional unimodal systems, limited to a single type of data, often fall short of capturing the complexity and richness of human communication and experience. In contrast, multimodal retrieval systems leverage the complementary nature of different data types to provide more accurate, context-aware, and user-centric search results. Text can provide specific details and context that images alone cannot convey. Conversely, images can instantly show concepts that might take longer to explain in words. Therefore, multimodal retrieval has wider applications in real world. For instance, consider you’ve previously visited a memorable location in New York City and captured a photo filled with landmarks and people. If you can’t recall the place’s name later, a multimodal system allows you to query “where is this place” along with the photo for identification. Healthcare is another domain where multimodal retrieval can be invaluable. Imagine a diagnostic support system that analyzes patients’ electronic health records, which contain a mix of textual data (like doctor’s notes), visual data (such as X-ray or MRI images), and even auditory data (like heart or lung sounds). A multimodal retrieval system can integrate these diverse data types to assist medical professionals in diagnosing complex conditions more accurately and swiftly.