Automatic Music Transcription (AMT) of piano music is a well-established study area. This paper addresses the broader problem of generating MIDI piano covers for pop music automatically. This is a more challenging task due to the diverse styles and instrumentation of pop music. In addition, there is a lack of large-scale dataset with paired and synchronized audio waveform and piano MIDI file. In this paper, we build a pipeline for data collection and synchronization using music information retrieval (MIR) techniques, resulting in a dataset of 3000 paired audio and piano MIDI file samples, with a total duration of 180 h. We propose two metrics for measuring the quality of synchronization, which is used to filter out poorly synchronized samples. Using a Transformer neural network, our model can generate piano cover song solely from the raw audio, without any input of external information. Our model outperforms recent similar works, Pop2Piano [2] and PiCoGen [25] in terms of melody chroma accuracy and subjective evaluation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Music2MIDI: Pop Music to MIDI Piano Cover Generation

  • Tin Yui Yip,
  • Chuck-jee Chau

摘要

Automatic Music Transcription (AMT) of piano music is a well-established study area. This paper addresses the broader problem of generating MIDI piano covers for pop music automatically. This is a more challenging task due to the diverse styles and instrumentation of pop music. In addition, there is a lack of large-scale dataset with paired and synchronized audio waveform and piano MIDI file. In this paper, we build a pipeline for data collection and synchronization using music information retrieval (MIR) techniques, resulting in a dataset of 3000 paired audio and piano MIDI file samples, with a total duration of 180 h. We propose two metrics for measuring the quality of synchronization, which is used to filter out poorly synchronized samples. Using a Transformer neural network, our model can generate piano cover song solely from the raw audio, without any input of external information. Our model outperforms recent similar works, Pop2Piano [2] and PiCoGen [25] in terms of melody chroma accuracy and subjective evaluation.