Big Data Mining Techniques and Their Optimisation in Financial Data Analysis
摘要
Big data mining techniques are crucial in modern financial data analysis. Large amounts of data are analysed in depth through advanced algorithms to extract the required information and transform and mine it appropriately to gain valuable insights. Big data mining objects include structured, semi-structured and unstructured data, and its core process covers data selection, preprocessing, transformation, mining and model evaluation, with each step providing a theoretical basis for systematic analysis. Stochastic decision trees are particularly important in financial data classification, which can increase the generalisation ability and robustness of the model, avoid overfitting and improve the prediction ability. Combining the methods of discriminant matrix and artificial weights can optimise the decision tree structure and improve the classification accuracy. Experiments show that the classification accuracy of the stochastic decision tree is relatively stable in different risk categories, but the accuracy of the high-risk category is lower, mainly due to the insufficient amount of high-risk data. By adding 500 high-risk data and adopting the stratified sampling method, the classification accuracy of the high-risk category was significantly improved while maintaining the high accuracy of other categories. Overall, the stratified sampling method significantly improves the classification model performance and provides reliable technical support for financial institutions’ risk assessment. The successful application of big data mining techniques in financial data analysis demonstrates its great potential in handling complex financial data and enhancing risk management capabilities.