Small, medium, and large organizations collect vast amounts of data with the expectation of using it to generate commercial value. Machine learning is a powerful tool for extracting valuable insights from this data and serves as a pivotal sales strategy for companies to maximize profits. This paper seeks to analyze sales data and discern patterns in sales among products that exhibit similarities, such as boxes and bags. In order to achieve this goal, was used unsupervised learning methods that allow the segmentation of groups, specifically Principal Component Analysis (PCA), k-means algorithms, and hierarchical clustering. PCA was used to identify correlated variables and find hidden patterns in the data, particularly pertaining to product families with similar sales. Elbow, Silhouette, and 30 indices methods were applied to determine the optimal number of clusters. Based on these results, it was determined the optimal number of clusters. Validation methods were employed to identify the clustering algorithm exhibiting the best performance. Stability measures evaluated the consistency of the clusters, while the cophenetic coefficient aided in determining the most effective data grouping method. After validation, the clustering algorithms were implemented. The results indicated that all clustering algorithms effectively segmented the data, with particular emphasis on the performance of the k-means algorithm. This study identified product groups with similar sales patterns and key products that impact the company’s global sales. Multivariate analysis provided a deeper understanding of sales dynamics, enabling the company to implement targeted marketing strategies and optimize resource allocation to boost bag and box sales in Portugal and other countries.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multivariate Analysis of Products Tipology Data - A Case Study

  • Nelson Costa,
  • Alzira Mota,
  • Inês Sousa

摘要

Small, medium, and large organizations collect vast amounts of data with the expectation of using it to generate commercial value. Machine learning is a powerful tool for extracting valuable insights from this data and serves as a pivotal sales strategy for companies to maximize profits. This paper seeks to analyze sales data and discern patterns in sales among products that exhibit similarities, such as boxes and bags. In order to achieve this goal, was used unsupervised learning methods that allow the segmentation of groups, specifically Principal Component Analysis (PCA), k-means algorithms, and hierarchical clustering. PCA was used to identify correlated variables and find hidden patterns in the data, particularly pertaining to product families with similar sales. Elbow, Silhouette, and 30 indices methods were applied to determine the optimal number of clusters. Based on these results, it was determined the optimal number of clusters. Validation methods were employed to identify the clustering algorithm exhibiting the best performance. Stability measures evaluated the consistency of the clusters, while the cophenetic coefficient aided in determining the most effective data grouping method. After validation, the clustering algorithms were implemented. The results indicated that all clustering algorithms effectively segmented the data, with particular emphasis on the performance of the k-means algorithm. This study identified product groups with similar sales patterns and key products that impact the company’s global sales. Multivariate analysis provided a deeper understanding of sales dynamics, enabling the company to implement targeted marketing strategies and optimize resource allocation to boost bag and box sales in Portugal and other countries.