Structure and design of multimodal dataset for automatic regex synthesis methods in Roman Urdu
摘要
Automatic regex synthesis involves generation of regular expressions from user-written natural language descriptions, example strings or both. Daily, countless regex generation queries are posted on online Q&A platforms such as StackOverflow (https://quora.com) and Quora (https://stackoverflow.com). Existing automatic regex synthesis methods demand concretely designed, multimodal datasets for optimal performance. Unfortunately, publicly available datasets even for resource-rich languages like English are often model-specific and incomplete, potentially hindering the efficiency and accurateness of regex synthesis methods. This issue is worsened for resource-poor languages such as Standard Urdu and Roman Urdu. In this paper, we present a novel, benchmark Roman Urdu dataset with 900 words and a novel Roman Urdu lexicon of 4225 words, annotated and labeled, to address the unmet needs of regex synthesis methods. Equipping these methods with a proficient dataset can lead to more fruitful regex generation.