<p>While machine learning offers powerful tools for predicting consumer behavior, the utility of marketing datasets is often undermined by two pervasive issues: substantial missing data and imbalanced class distributions across customer groups. These challenges are especially acute in binary classification problems common in marketing, such as campaign response prediction (accept/reject) and purchase decision modeling (buy/not buy), where missing values and class imbalance can markedly degrade model performance. This study systematically evaluates preprocessing strategies to address both issues simultaneously, using 32,860 real-world consumer records from a financial services firm. We evaluated our framework using five machine learning models: K-Nearest Neighbors, Decision Trees, Random Forest, Multi-Layer Perceptron, and AdaBoost classifiers. The findings advance marketing analytics by demonstrating how tailored data preparation strengthens the reliability of AI-driven consumer insights, particularly for binary classification tasks with imbalanced class distributions (e.g., campaign responses and purchase conversions). In doing so, the proposed approach helps bridge the gap between theoretical machine learning methods and their practical application in marketing contexts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A machine learning framework for missing and imbalanced data in marketing analytics

  • Chi Zhang,
  • Wenkai Zhou,
  • Xiaowei Zhao,
  • Shengfeng Yang

摘要

While machine learning offers powerful tools for predicting consumer behavior, the utility of marketing datasets is often undermined by two pervasive issues: substantial missing data and imbalanced class distributions across customer groups. These challenges are especially acute in binary classification problems common in marketing, such as campaign response prediction (accept/reject) and purchase decision modeling (buy/not buy), where missing values and class imbalance can markedly degrade model performance. This study systematically evaluates preprocessing strategies to address both issues simultaneously, using 32,860 real-world consumer records from a financial services firm. We evaluated our framework using five machine learning models: K-Nearest Neighbors, Decision Trees, Random Forest, Multi-Layer Perceptron, and AdaBoost classifiers. The findings advance marketing analytics by demonstrating how tailored data preparation strengthens the reliability of AI-driven consumer insights, particularly for binary classification tasks with imbalanced class distributions (e.g., campaign responses and purchase conversions). In doing so, the proposed approach helps bridge the gap between theoretical machine learning methods and their practical application in marketing contexts.