Analysis of Indonesian Multiword Expressions: Linguistic vs Data-Driven Approach
摘要
Tagging systems developed using a data-driven approach are often considered superior to those produced using a linguistic approach [Brill (A Simple Rule-Based Part of Speech Tagger. Applied Natural Language Processing Conference, 1992, p.152)]. The creation of dictionaries and grammars (resources typically used in a linguistic approach) is considered costly compared to the creation of a training corpus (a resource typically used in a data-driven approach) [Silberztein (Formalizing Natural Languages: The NooJ Approach, 2016, p.22)]. In this contribution, I argue that such a view needs to be reconsidered. Focusing on MWE, I will show that some data-driven systems which rely on training corpora may produce inaccurate results, leading to incorrect automatic POS tagging, syntactic parsing and machine translation. I also show that such errors can be prevented using dictionaries and grammars for systems developed using a linguistic approach, which is principally in line with Silberztein’s (Formalizing Natural Languages: The NooJ Approach, 2016) view.