Tabular data, known for its interpretability and processing efficiency, is widely used in various applications such as disease risk prediction, credit scoring, and advertising. However, existing methods are often restricted by the class imbalance problem, leading to biased decision-making and poor performance in identifying the minority class. To address this challenge, this study proposes a novel dual-layer compensation framework that combines data-level augmentation and model-level enhancements to improve classification performance. First, synthetic minority over-sampling technique (SMOTE) is employed to balance the dataset distribution by artificially synthesizing minority class samples. Then, a Transformer-based framework, CATransformer, is designed with Class-Aware Attention (CAA) and Gated Linear Unit (GLU). Specifically, to improve the model’s ability to differentiate between semantic features across various categories and to compensate for the limited diversity of synthetic samples, we design a CAA module. This module learns trainable bias vectors for each category, which are integrated into the multi-head attention scores, allowing the model to dynamically focus on key semantic subspaces relevant to different categories. Additionally, we optimize the feedforward neural network (FFNN) structure within the Transformer by proposing an improved variant of the GLU. Finally, we conducted comprehensive experiments on multiple public datasets to evaluate the proposed method. The proposed method obtains comparable performance to the best baseline. Meanwhile, compared with the Transformer-based model FT-Transformer, our approach achieves improvements of 1.5%, 4.7%, and 2.9% in Precision, F1-score, and G-mean, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CATransformer: A Class-Aware Transformer Within a Dual-Layer Framework for Imbalanced Tabular Data

  • Yifan Wang,
  • Haiyong Shi,
  • Zhengang Guo,
  • Hanxi Zou,
  • Bingyi Liu

摘要

Tabular data, known for its interpretability and processing efficiency, is widely used in various applications such as disease risk prediction, credit scoring, and advertising. However, existing methods are often restricted by the class imbalance problem, leading to biased decision-making and poor performance in identifying the minority class. To address this challenge, this study proposes a novel dual-layer compensation framework that combines data-level augmentation and model-level enhancements to improve classification performance. First, synthetic minority over-sampling technique (SMOTE) is employed to balance the dataset distribution by artificially synthesizing minority class samples. Then, a Transformer-based framework, CATransformer, is designed with Class-Aware Attention (CAA) and Gated Linear Unit (GLU). Specifically, to improve the model’s ability to differentiate between semantic features across various categories and to compensate for the limited diversity of synthetic samples, we design a CAA module. This module learns trainable bias vectors for each category, which are integrated into the multi-head attention scores, allowing the model to dynamically focus on key semantic subspaces relevant to different categories. Additionally, we optimize the feedforward neural network (FFNN) structure within the Transformer by proposing an improved variant of the GLU. Finally, we conducted comprehensive experiments on multiple public datasets to evaluate the proposed method. The proposed method obtains comparable performance to the best baseline. Meanwhile, compared with the Transformer-based model FT-Transformer, our approach achieves improvements of 1.5%, 4.7%, and 2.9% in Precision, F1-score, and G-mean, respectively.