错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Textual Data

  • Taylor Arnold,
  • Lauren Tilton

摘要

We provide an introduction to running and applying natural language processing (NLP) techniques to a corpus of textual documents. NLP methods include tokenization, lemmatization, and part-of-speech tagging. Using the automatically tagged data, techniques such as TF-IDF, document distance, and dimensionality reduction are applied to a collection of documents.