MakedonASRDataset - A Dataset for Speech Recognition in the Macedonian Language
摘要
Using dataset analysis as a research method is becoming more popular among many researchers with diverse data collection and analysis backgrounds. This paper provides the first publicly available dataset consisting of audio segments and appropriate textual transcription in the Macedonian language. It is appropriately preprocessed and prepared for direct utilization in the automatic speech recognition pipelines. The dataset was created by students at the Faculty of Computer Science and Engineering as part of the elective course, ‘Digital Libraries’, with the audio segments sourced from a YouTube channel.