错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MakedonASRDataset - A Dataset for Speech Recognition in the Macedonian Language

  • Martin Mishev,
  • Blagica Penkova,
  • Maja Mitreska,
  • Magdalena Kostoska,
  • Ana Todorovska,
  • Monika Simjanoska,
  • Kostadin Mishev

摘要

Using dataset analysis as a research method is becoming more popular among many researchers with diverse data collection and analysis backgrounds. This paper provides the first publicly available dataset consisting of audio segments and appropriate textual transcription in the Macedonian language. It is appropriately preprocessed and prepared for direct utilization in the automatic speech recognition pipelines. The dataset was created by students at the Faculty of Computer Science and Engineering as part of the elective course, ‘Digital Libraries’, with the audio segments sourced from a YouTube channel.