Construction of Data Stream Classification Model Based on Machine Learning Algorithm
摘要
Traditional data categorization methods, such as decision trees and ANN (Artificial Neural Network), are hard to apply when dealing with data streams, and even fail because they focus only on static data extraction. With the advent of the big data era, IoT technology poses new challenges to traditional data mining techniques. Data streams are characterized by three features: real-time, large data volumes, and dynamic changes over time. For the characteristics of real-time and large data volume, there are some relatively mature algorithms that can process a large amount of data in real time, thus affecting the classification model by improving the speed of the classification model. This paper is based on the above problems, discusses the construction of data flow classification model based on machine learning (ML) algorithms in the context of information technology, and takes the text data classification algorithm based on HMM (Hidden Markov Model) algorithm as an example, and finally experimentally verifies the classification time of the HMM algorithm is faster than that of the Hidden Markov Model (HMM) algorithm when the number of dataset instances is less than in the 4 × 106 the classification time of HMM algorithm is less than that of Bagging algorithm.