This work presents a novel approach to automatic music generation using the Apple Loops library of music fragments. Each loop fragment is a pre-recorded sound of one instrument lasting a few seconds. By leveraging the capabilities of generative adversarial networks (GANs) and deep learning techniques, our automatic music generator creates realistic Hip-Hop compositions by arranging loops in a structured manner. Results demonstrate the feasibility of the proposed approach, proving the concept and showing promising outcomes in terms of artistic quality. An ablation study was used to evaluate the role of individual model components. It is found that the representation of loops as vectors using audio waveform compression yields better results because of its compatibility with the backpropagation algorithm. Future work will expand the method to control over the emotional flavor of the generated music. This feature, if used in real time, can add a new nonverbal modality to the human-computer social interaction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Music Generation Using a Library of Music Fragments

  • Maxim S. Gavrilov,
  • Alexei V. Samsonovich

摘要

This work presents a novel approach to automatic music generation using the Apple Loops library of music fragments. Each loop fragment is a pre-recorded sound of one instrument lasting a few seconds. By leveraging the capabilities of generative adversarial networks (GANs) and deep learning techniques, our automatic music generator creates realistic Hip-Hop compositions by arranging loops in a structured manner. Results demonstrate the feasibility of the proposed approach, proving the concept and showing promising outcomes in terms of artistic quality. An ablation study was used to evaluate the role of individual model components. It is found that the representation of loops as vectors using audio waveform compression yields better results because of its compatibility with the backpropagation algorithm. Future work will expand the method to control over the emotional flavor of the generated music. This feature, if used in real time, can add a new nonverbal modality to the human-computer social interaction.