Dataset Balancing
摘要
Many classification data mining studies involve highly skewed data, such as bankruptcy and medical issues (both of which are hoped to be rare). This can lead to statistical issues. Methods for dataset balancing are discussed and demonstrated on four different bankruptcy data files. They are also applied to a credit card fraud detection dataset.