Addressing the data gap: building a parallel corpus for Kashmiri language
摘要
This paper marks a significant step forward in language technology for low-resource languages by developing the first parallel corpus for the Kashmiri language, which previously lacked substantial digital resources. We compiled and refined approximately 30,000 sentence pairs through innovative data collection and processing techniques, establishing a high-quality corpus. Leveraging this corpus, we built a Neural Machine Translation (NMT) model, demonstrating its effectiveness with comprehensive performance metrics. Our findings not only showcase the potential for enhancing NMT systems for languages like Kashmiri but also lay the groundwork for future research in linguistic technology development without the need for extensive external resources.