Methods and Technologies for Streaming Primary Processing and Analysis of Big Data from Multi-Assortment Production for Predicting Polymeric Film Quality
摘要
Methods and algorithms for streaming processing and analysis of big production data have been proposed, which allow loading arrays of parameter values of industrial production of polymeric films collected from various sources in real time, carrying out their primary processing, storing in a database of production parameters, statistical analysis and constructing predictive models to estimate consumer characteristics of products. Primary processing includes the stages of converting data from service codes into real values, filtering and structuring data by production stages and parameter types. Mathematical models of physical processes in production units are used to calculate quality indicators of polymeric material at stages that are not measured, but are important for product quality control. The algorithm of statistical analysis consists in checking the conformity of the distribution of production data to the normal distribution and estimating correlations to select the control actions that most strongly affect the consumer characteristics of products. Methods of mathematical statistics (linear multiple regression) and data mining (neural network, boosting of decision trees) are used to construct predictive models, depending on the type of distribution law, volume of production data, requirements for accuracy and cost-effectiveness of the forecast. Algorithms for processing and analyzing big production data have been implemented in the form of software modules for the software package of big data mining to predict and control the quality of polymeric films. Testing of the software, according to the data of the production of pharmaceutical and food rigid packaging films at the Russian factory, has confirmed its operability.