A Study on Synthesizing Expressive Violin Performances: Approaches and Comparisons
摘要
Expressive music synthesis (EMS) for violin performance is a challenging task due to the disagreement among music performers in the interpretation of expressive musical terms (EMTs), the paucity of labeled recordings, and the limited generalization ability of the synthesis model. These challenges create trade-offs between model effectiveness, diversity of generated results, and controllability of the synthesis system, making it essential to conduct a comparative study on EMS model design. In this paper, we discuss two violin EMS approaches. First, the end-to-end approach is a modification of a state-of-the-art text-to-speech neural network. Second, the parameter-controlled approach is based on a simple parameter sampling process that can render note lengths and other parameters compatible with MIDI-DDSP. We systematically study the two approaches through both objective and subjective experiments and discuss several key issues of EMS based on the results.