Syntactics Processing
摘要
Computational syntactics processing is a fundamental technique in NLU. It normally serves as a preprocessing method to transform natural language into structured and normalized texts, yielding syntactic features for downstream task learning. This chapter proposes a systematic review of low-level syntactic processing techniques, namely: microtext normalization, sentence boundary disambiguation, POS tagging, text chunking, and lemmatization. In particular, we summarize and categorize widely used methods in the aforementioned syntactic analysis tasks, investigate the challenges, and yield possible research directions to overcome the challenges in future work.