Leveraging Multi-modal Data
摘要
This chapter examines how multi-modal data like text, image, audio, and videos can enhance recommendation systems. It introduces core integration strategies (early, late, and hybrid fusion) and contrasts them with emerging multi-modal large language models (LLMs) in terms of architecture, training, and use cases.