We introduce a novel method for combining two deep neural networks (DNNs) using a Wasserstein GAN (WGAN)-based adapter, aimed at facilitating efficient transfer learning. Traditional transfer learning techniques often require significant computational resources and large amounts of data to achieve high performance. Our approach leverages pre-trained models to mitigate these demands, focusing on transforming intrinsic data representations between specified layers of deep networks. Initially, we identify the layers within the pre-trained networks, where the adapter should be connected. The WGAN-based adapter is then trained to approximate the transformation between these layers, significantly reducing the number of layers and model parameters compared to the original networks while maintaining accuracy. We apply our method to the task of speech audio classification. Our experiments demonstrate that the proposed method reduces the need for extensive training data and computational resources, offering an efficient and scalable solution for combining DNNs. This approach enhances model performance and robustness, providing significant improvements in various tasks such as image classification, speech recognition, and natural language processing. Our results suggest that the WGAN-based adapter holds potential as an efficient tool for merging the knowledge from machine learning and artificial intelligence models in broader applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Wasserstein GAN-Based Adapter for Deep Neural Networks Merging

  • Miron M. Leonov,
  • Artem A. Soroka,
  • Alexander G. Trofimov

摘要

We introduce a novel method for combining two deep neural networks (DNNs) using a Wasserstein GAN (WGAN)-based adapter, aimed at facilitating efficient transfer learning. Traditional transfer learning techniques often require significant computational resources and large amounts of data to achieve high performance. Our approach leverages pre-trained models to mitigate these demands, focusing on transforming intrinsic data representations between specified layers of deep networks. Initially, we identify the layers within the pre-trained networks, where the adapter should be connected. The WGAN-based adapter is then trained to approximate the transformation between these layers, significantly reducing the number of layers and model parameters compared to the original networks while maintaining accuracy. We apply our method to the task of speech audio classification. Our experiments demonstrate that the proposed method reduces the need for extensive training data and computational resources, offering an efficient and scalable solution for combining DNNs. This approach enhances model performance and robustness, providing significant improvements in various tasks such as image classification, speech recognition, and natural language processing. Our results suggest that the WGAN-based adapter holds potential as an efficient tool for merging the knowledge from machine learning and artificial intelligence models in broader applications.