DuFCALF: Instilling Sentience in Computerized Song Analysis
摘要
Music recommendation systems have evolved significantly in the past couple of years and have become extremely popular with the advancement of Artificial Intelligence (AI). Such systems categorize songs based on disparate perspectives, like genre, artist, tempo, etc. However, there have been fewer developments in the purview of song categorization based on feelings. Songs amplify the mood of a person and often, listeners choose the type of song they want to listen to based on their mood. Hence, modeling the emotional content of songs is crucial. However, this is a challenging affair because songs embody multiple instruments and vocals in a single instance. Each of them is often modulated differently for a better experience for the listeners as well. This is very common for present-day songs and many songs of different emotions often have very similar instruments and chord progressions which further enhances the challenge. In this paper, a system is presented to distinguish song clips based on their emotion. The clips were parameterized using two features which were fed to disparate deep networks. Thereafter, a dual feature cross architecture late fusion (DuFCALF) strategy was used to distinguish the moods. Experiments were performed with multitudinous sections of songs to test its efficacy for limited data and the sadness of songs was captured with over 80% accuracy. An overall increase of 4.43% in distinction of the emotions was obtained using DuFCALF over the best-performing baseline system.