Full-Page Music Symbols Recognition: State-of-the-Art Deep Model Comparison for Handwritten and Printed Music Scores
摘要
The localization and classification of musical symbols on scanned or digital music scores pose significant challenges in the process of Optical Music Recognition. For instance, similar musical symbol classes and a large number of overlapping tiny musical symbols within high-resolution music scores appear in musical scores. Recently, deep learning-based techniques show promising results in addressing these challenges by leveraging object detection models. However, unclear directions in training and evaluation approaches, such as inconsistency between usage of full-page or cropped images, handling image scores at full-page level in high-resolution, reporting results on only specific object classes, missing comprehensive analysis with recent state-of-the-art object detection methods, cause a lack of benchmarking and of analyzing the impact of proposed methods in music object recognition. To address these issues, we perform intensive analysis with recent object detection models, exploring effective ways of handling high-resolution images on existing benchmarks. Our goal is to bridge the gap between object detection models designed for common objects and relatively small images compared to music scores, and the unique challenges of music score recognition in terms of object size and resolution. We achieve state-of-the-art results across mAP and Weighted mAP on two challenging datasets, namely DeepScoresV2 and the MUSCIMA++ datasets, by demonstrating the effectiveness of this approach in both printed and handwritten music scores.