Single Modality to Multi-modality: The Evolutionary Trajectory of Artificial Intelligence in Integrating Diverse Data Streams for Enhanced Cognitive Capabilities
摘要
The evolution of artificial intelligence (AI) towards multi-modality represents a significant advance in the field by integrating multiple sensory inputs to enhance machine understanding and affected communication. This research provides a broad overview of these processes, tracing the history of the development of artificial intelligence from its early stages to the rise of machine learning. It delves into the foundations of multi-modal AI, covering its concepts, values, and early stages. Core enabling technologies, including neural networks, transformers, tracking mechanisms, multi-modal fusion technology, and cross-modal learning, are examined. Multi-modal AI systems, such as vision-language models (e.g., CLIP and DALL-E), audio-visual models, text and image synthesis models, and diversity of reference standards, are explored along with their applications in natural language processing, computer vision, speech and visualisation, robotics, and healthcare. Developments in multi-disciplinary interaction are discussed with a focus on user interaction, human-computer interaction, and augmented and virtual reality. The impact of multi-modal intelligence on creative industries such as art, design, film, animation, music, and games is also examined. Ethical and social aspects such as impartiality, fairness, privacy, transparency, and efficiency are examined. Future directions and innovations in multi-modal AI are explored, focusing on advances in learning algorithms, integration with new technologies, achievements over time, and personal development and adaptation. The study aims to provide a better understanding of change and the problems that need to be solved.