<p>Automatic music generation plays a crucial role in generating creative compositions autonomously, facilitating applications in various fields, including entertainment and education. The challenges faced by existing approaches include capturing temporal dependencies, ensuring harmonic coherence, and providing scalability and flexibility. In this paper, a novel Multi-Module Neural Network (MNN) is proposed to solve these challenges. Temporal dependencies in the chord progression are modeled using a Temporal Graph Dilated Convolution (TGDC) layer. Inside the Rhythm Generator, rhythmic patterns are produced by a MobileNet Temporal Convolutional Network (MobileNetTCN), while the Pitch Generator utilizes the Sparse Transformer for the coherent generation of pitch sequences aligned to the rhythm. The outputs of these modules are consolidated to generate a coherent melody; their timing, velocity, and harmonic coherence are post-processed. Lastly, the produced music is converted to audio or MIDI form for playback. Experimental evaluation validates that the proposed framework achieves 98.98% Accuracy in predicting musical events, including note timings, chord structures, and harmonic transitions, along with 97.83% Precision, 96.92% Recall, 97.37% F1 score, a Bar Rhythm Analysis (BRA) of 98.52%, and a Chord Tone Ratio (CTR) of 0.983. These findings highlight the framework's potential to set new standards in the field of automatic music generation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Music Generation with Multi-module Neural Networks for Chord, Rhythm, and Pitch Modeling

  • Pengzhan Qin

摘要

Automatic music generation plays a crucial role in generating creative compositions autonomously, facilitating applications in various fields, including entertainment and education. The challenges faced by existing approaches include capturing temporal dependencies, ensuring harmonic coherence, and providing scalability and flexibility. In this paper, a novel Multi-Module Neural Network (MNN) is proposed to solve these challenges. Temporal dependencies in the chord progression are modeled using a Temporal Graph Dilated Convolution (TGDC) layer. Inside the Rhythm Generator, rhythmic patterns are produced by a MobileNet Temporal Convolutional Network (MobileNetTCN), while the Pitch Generator utilizes the Sparse Transformer for the coherent generation of pitch sequences aligned to the rhythm. The outputs of these modules are consolidated to generate a coherent melody; their timing, velocity, and harmonic coherence are post-processed. Lastly, the produced music is converted to audio or MIDI form for playback. Experimental evaluation validates that the proposed framework achieves 98.98% Accuracy in predicting musical events, including note timings, chord structures, and harmonic transitions, along with 97.83% Precision, 96.92% Recall, 97.37% F1 score, a Bar Rhythm Analysis (BRA) of 98.52%, and a Chord Tone Ratio (CTR) of 0.983. These findings highlight the framework's potential to set new standards in the field of automatic music generation.