<p>Credit risk assessment is a critical task in the lending industry, where private credit data from users are collected and processed to evaluate their ability to repay loans. With the development of AI technologies, financial institutions are increasingly adopting AI models to improve evaluation efficiency and accuracy. However, due to the existence of data silos between institutions, the sharing of data across institutions is difficult, limiting the further enhancement of model performance. Multi-institutional collaborative learning can effectively integrate data resources; however, the need for institutions to build personalized models while protecting privacy contradicts the data-sharing concept. This paper introduces an efficient collaborative learning framework for multi-institution scenarios to tackle this issue, incorporating knowledge filtering and data sharing mechanisms. The framework consists of the following key steps: (1) Feature analysis of data from each institution is conducted to build a unified feature space, thereby alleviating data heterogeneity. (2) Each institution generates synthetic data based on local data using CTGAN, and the synthetic data is filtered through a designed data filtering algorithm to obtain a streamlined pseudo-dataset. (3) Institutions share the filtered pseudo-dataset as needed and, through a phased data fusion mechanism, gradually integrate shared data from other institutions with local datasets to form an enhanced local dataset, thereby improving model performance. Experimental results show that the proposed framework performs significantly well under independent and identically distributed (iid) and non-independent and identically distributed (non-iid) scenarios. For the best-performing model, our method improves the accuracy by 8% and the recall rate by 35% after only one round of communication, thoroughly verifying that the proposed framework can provide each institution with high-performance personalized models while preserving privacy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An efficient multi-institution collaborative learning framework based on knowledge filtering and data sharing

  • Zhijian Li,
  • Guanxin Chen,
  • Yipeng Liu,
  • Jianting Yuan

摘要

Credit risk assessment is a critical task in the lending industry, where private credit data from users are collected and processed to evaluate their ability to repay loans. With the development of AI technologies, financial institutions are increasingly adopting AI models to improve evaluation efficiency and accuracy. However, due to the existence of data silos between institutions, the sharing of data across institutions is difficult, limiting the further enhancement of model performance. Multi-institutional collaborative learning can effectively integrate data resources; however, the need for institutions to build personalized models while protecting privacy contradicts the data-sharing concept. This paper introduces an efficient collaborative learning framework for multi-institution scenarios to tackle this issue, incorporating knowledge filtering and data sharing mechanisms. The framework consists of the following key steps: (1) Feature analysis of data from each institution is conducted to build a unified feature space, thereby alleviating data heterogeneity. (2) Each institution generates synthetic data based on local data using CTGAN, and the synthetic data is filtered through a designed data filtering algorithm to obtain a streamlined pseudo-dataset. (3) Institutions share the filtered pseudo-dataset as needed and, through a phased data fusion mechanism, gradually integrate shared data from other institutions with local datasets to form an enhanced local dataset, thereby improving model performance. Experimental results show that the proposed framework performs significantly well under independent and identically distributed (iid) and non-independent and identically distributed (non-iid) scenarios. For the best-performing model, our method improves the accuracy by 8% and the recall rate by 35% after only one round of communication, thoroughly verifying that the proposed framework can provide each institution with high-performance personalized models while preserving privacy.