<p>Deep learning in speech-related tasks has made significant strides, primarily driven by innovations such as transformer architectures and increased computational power. However, progress has been uneven across languages, with a strong bias toward English due to the availability of large, high-quality datasets. In contrast, Farsi (Persian) remains underrepresented, hindered by a lack of comparable resources. To bridge this gap, we introduce the first publicly available word-level Farsi speech corpus, PAZHVAK. The dataset contains 4018 unique words spoken by 61 participants (38 male, 23 female), totaling 88,535 recordings and over 56&#xa0;h of audio. The recordings were captured in diverse environments and with varying intonations. Each clip, lasting between 0.5 and 7&#xa0;s, is sampled at 16 kHz in mono. This corpus provides a clean and comprehensive foundation for training and evaluating speech-related models in Farsi. It is freely accessible at <a href="https://hormozgan.ac.ir/home/index/33/91/2164">PAZHVAK</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PAZHVAK: a Word-Level Farsi Speech Corpus by University of Hormozgan

  • Mohammad Azim Saraji,
  • Abdullah Khalili,
  • Ahmad Hatam

摘要

Deep learning in speech-related tasks has made significant strides, primarily driven by innovations such as transformer architectures and increased computational power. However, progress has been uneven across languages, with a strong bias toward English due to the availability of large, high-quality datasets. In contrast, Farsi (Persian) remains underrepresented, hindered by a lack of comparable resources. To bridge this gap, we introduce the first publicly available word-level Farsi speech corpus, PAZHVAK. The dataset contains 4018 unique words spoken by 61 participants (38 male, 23 female), totaling 88,535 recordings and over 56 h of audio. The recordings were captured in diverse environments and with varying intonations. Each clip, lasting between 0.5 and 7 s, is sampled at 16 kHz in mono. This corpus provides a clean and comprehensive foundation for training and evaluating speech-related models in Farsi. It is freely accessible at PAZHVAK.