Amazlem: The First Amazigh Lemmatizer
摘要
Natural language processing has become the center of research, not only, in rich languages but also in low resourced ones. In this perspective and to enrich Amazigh as an under-resourced language, we present in this paper, Amazlem, the first lemmatizing system for Amazigh language. This system takes as input, sentences in Amazigh language and outputs the lemma of each word in the sentences. The approach considered in the realization of this system is based on the rules from the Amazigh language grammar. These rules include the formation of nouns and verbs in Amazigh language. Starting from the word even without knowing its morphological syntax, the system returns the lemma of the word by matching it in a dictionary. Amazlem uses two lemmatization algorithms, one for nouns and the other for verbs. Due to the absence of a labeled corpus of lemma, we validate this approach by labeling a dataset of 10000 words from an existing Amazigh corpus for part-of-speech tagging and the results reaches 85, 5%.