Unsupervised ML with Text Data
摘要
This chapter builds on the techniques introduced in the previous Chaps. 12 and 13 . Specifically, we will demonstrate how unsupervised pattern recognition approaches can be applied to text data to answer a research question. Because pre-processing and “knowing” your data are especially important when using unsupervised approaches with text, we will introduce additional techniques with an emphasis on exploratory data analysis tools such as token frequency analysis, basic text analytics, and n-gram analysis to explore large text-based datasets before using this information in support of the application of unsupervised natural language processing (NLP) techniques.