Autoencoder as Feature Extraction Technique for Financial Distress Classification
摘要
Financial statements are typical financial distress identification data for the enterprise. However, nowadays, the valuable data source characterizing enterprise could be expanded, including data from legal events, macro, industry, government register center, etc. This data creates valuable information, which could lead to more accurate financial distress classification model creation. On the other hand, the new data source involvement expands the dimensional space of features and increases the data sparsity. In order to reduce dimensions and have maximum information retention from the initial data space is used feature extraction techniques. This study uses an autoencoder as a nonlinear feature extraction method. Moreover, we compared several structure composition strategies for autoencoders: 1) all data compress; 2) union of the several autoencoders (i.e. data compress of each data type separately and the union of these separate autoencoders). After implementing different autoencoder strategies, eight machine-learning models for financial distress classification were used. The results demonstrated that features retrieved from the union data source strategy outperform the features extracted all at once. These findings create a novelty of autoencoder usage as a feature extraction technique for financial distress key feature’s identification and better financial distress issue classification.