Representation and Analysis of Dynamics for Automated Music Assessment in Hindustani Vocal Music
摘要
Automatic music assessment systems rely on musically relevant representations of music dimensions to provide accurate and meaningful feedback to music learners on these dimensions. Dynamics is one such fundamental dimension of music performance, in addition to melody and rhythm, contributing to the expressivity and musical expression. However, unlike melody and rhythm, music representations that encode dynamics have received limited attention. Dynamics are often extracted using objective loudness measures from an audio music piece, but those measures need to be translated into suitable abstract musically meaningful representations to provide accurate feedback and assessment for effective music learning. While Western popular music has a systematic methodology to represent loudness, it is not the case with music cultures that are learned largely through the process of imitation, such as Hindustani vocal music, where the encoding of dynamics is implicit. The absence of a well defined, accurate, descriptive framework to represent and describe dynamics along with melody, rhythm, and other ornamentation leads to challenges in automatic music learning and assessment for Hindustani music. Srinivasamurthy and Chordia [9] developed a machine-readable unified framework to encode the bhatkhande symbolic representation of Hindustani music pieces. We propose a methodology to extract loudness descriptors from the audio that can further be encoded in the extended framework to map to dynamics. We perform data analysis on an example rendition followed by extended analysis on the saraga dataset showing the dynamics variation in the performances, with a goal to build accurate, descriptive scores of music recordings of Hindustani music that can encode melody, rhythm, and dynamics, and further be used to provide feedback to music learners.