An Accidental Benchmark: The History, Contingent Power, and Lasting Traces of the GTZAN Dataset
摘要
In 2002, George Tzanetakis presented a paper on how researchers could automatically classify musical genre from audio signals. Claiming that his model worked as well as human classifiers, Tzanetakis made his dataset available to anyone who asked for it. Ten years later, a systematic review found that this dataset had circulated on a massive scale — nearly 25% of papers on Music Genre Recognition (MGR) used the so-called GTZAN dataset in their research. Yet an analysis of the dataset revealed significant problems: repetitions, overrepresentations, files distorted to the point of corruption, with few researchers indicating that they ever listened to the musical files within it. These warnings went unheeded: the GTZAN dataset remains the most widely used dataset for MGR today. In this paper, I examine the GTZAN dataset from a historical and musicological perspective. I trace the dataset’s introduction into the Music Information Retrieval (MIR) community, and show how MIR researchers’ tendency to view the digital musical object as a set of extracted statistical features, and musical genre as a static, query-able combination of those features, created so-called ground truths about music that remain embedded in our present-day digital infrastructures. I argue that tracing this dataset’s history and ascendence to benchmark status can recover the ground-truthing process used by early MIR researchers. In addition, this history provides context for the music industry’s shift towards descriptive audio tagging and context-based recommendations.