Credit risk prediction and heterogeneity analysis for SMEs based on large language models and multimodal data fusion
摘要
Small and medium-sized enterprises (SMEs) play a crucial role in the national economy, but they face significant challenges in credit risk management. This is reflected in the obvious defects of relying solely on Finance metrics and the limitations of traditional large language models (LLMs) in processing the interaction between numerical and literal features. Specifically, the issue of finance metric fraud has caused critical damage to the reliability of metrics, and the neglect of non-financial factors such as administer competency and news public opinion makes it impossible to systematically quantize key risks including reputation exposure and external shocks. Meanwhile, the limitations of traditional language models like BERT lie in their inability to grab the integration relationship between numerical traits and literal traits. During risk forecast, they can only mechanically assemble features of different modalities through the idea of Ensemble Learning. Based on the above practical issues, this study proposes a credit risk forecast framework for SMEs based on large language models. This framework integrates multi-modality traits such as Finance metrics, non-Finance metrics, and unstructured text, and automatically learns the deep interaction relations among different modality traits to conduct risk forecast. It innovatively adopts Low-Rank Adaptation (LoRA) to fine tune the pre-trained model, improving forecast accuracy while reducing calculate cost. The empirical outcome shows that multi-modality data Fusion significantly improves model performance: the F1 score of the LLM-LoRA model is 0.829, which is 50.3% higher than the model using only Finance data. Meanwhile, to explore the contribution of different factors to credit risk forecast, this paper conducts trait heterogeneity analysis and puts forward rational recommendations for different principals based on the outcome.