错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-Media Intelligence

  • Yunhe Pan,
  • Yueting Zhuang,
  • Fei Wu,
  • Yi Yang,
  • Jun Xiao,
  • Tiejun Huang,
  • Wenwu Zhu,
  • Shiliang Zhang

摘要

The rapid growth of the internet and data explosion have driven media integration, making cross-media a key expressive form and positioning cross-media computing as a vital research area in multimedia. This chapter focuses on core challenges and recent advances in cross-media intelligence, covering three main areas: cross-media association, cross-modal reasoning, and large cross-modal models. Cross-media association explores alignment and understanding of multimodal content using methods like CCA, topic models, adversarial learning, and graph neural networks. Cross-modal reasoning involves tasks such as image captioning, speech synthesis, and evidence inference. Recently, large models like CLIP and GPT-4, along with domain-specific variants, have significantly advanced applications in fields like chemistry. Cross-media intelligence is also driving transformation in education, healthcare, and public safety. Future research will emphasize algorithm optimization, knowledge-driven approaches, and model interpretability to advance intelligent computing paradigms.