CoCoFormer: A Controllable Feature-Rich Polyphonic Music Generation Method
摘要
This paper explores the modeling method of polyphonic music sequence. Due to the great potential of Transformer models in music generation, controllable music generation is receiving more attention. In the task of polyphonic music, current controllable generation research focuses on controlling the generation of chords but lacks precise adjustment for the controllable generation of choral music textures. This paper proposes a Condition Choir Transformer (CoCoFormer) which controls the model’s output by controlling the input of the chord and rhythm at a fine-grained level. This paper’s self-supervised method improves the loss function and performs joint training through conditional control input and unconditional input training. This paper adds an adversarial training method to alleviate the lack of diversity in generated samples caused by teacher-forcing training. CoCoFormer enhances model performance with explicit and implicit inputs to chords and rhythms. In this paper, the experiments show that CoCoFormer has reached much better performance when compared to existing approaches. Based on specifying the polyphonic music texture, the same melody can also be generated in various ways.