k-Nearest Neighbors
摘要
Given a metric space and a labeled dataset within it, we discuss several algorithms based on the concept of k-nearest neighbors. These include the k-NN classifier with majority vote and the k-NN regressor with arithmetic mean. The effect of overfitting is illustrated via several examples. We introduce some preprocessing methods and then generalize the initially mentioned setting of metric spaces to distance measures in order to include cosine similarity and cosine distance into our theory. As examples, we discuss text mining, product reviews, and handwriting recognition.