As startups are significant for financial expansion and technological advancement, it is critical to analyse them in order to identify patterns within the sector and potential areas for funding. The analysis of its initial stages data using different approaches to clustering is the primary thrust of this work. The algorithms employed include KMeans, Agglomerative Clustering, and Gaussian Mixture Models (GMM), with the incorporation of Principal Component Analysis (PCA) for dimensionality reduction. The ideal quantity of groups (k) for K-means and Agglomerative Filtering is found using the elbow approach. The computational methods are then used, and metric values are used to assess the clustering outcomes. The results indicate that KMeans achieved a Silhouette Score of 0.30, a Calinski-Harabasz Index of 682.37, and a Davies-Bouldin Index of 0.73. GMM yielded a Silhouette Score of 0.29, a Calinski-Harabasz Index of 2.277, and a Davies-Bouldin Index of 365.87. Agglomerative Clustering produced a Silhouette Score of 0.90, a Calinski-Harabasz Index of 331.54, and a Davies-Bouldin Index of 0.87. These results imply that, out of the three categories of algorithms, GMM operated the best. It achieved the highest Silhouette Score and relatively lower Davies-Bouldin Index, indicating better-defined and more separated clusters. All things considered, this work offers insightful information about how methodologies for clustering can be applied to startup data analysis, emphasising the significance of choosing suitable algorithms and assessment metrics in accordance with the features of the dataset and the clustering objectives.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Clustering Analysis of High-Tech Startups Using Machine Learningh

  • Devki Pradeep Deshpande,
  • Ankush D. Sawarkar,
  • Megha V. Jonnalagedda,
  • Shrinivas Khedkar,
  • Lal Singh

摘要

As startups are significant for financial expansion and technological advancement, it is critical to analyse them in order to identify patterns within the sector and potential areas for funding. The analysis of its initial stages data using different approaches to clustering is the primary thrust of this work. The algorithms employed include KMeans, Agglomerative Clustering, and Gaussian Mixture Models (GMM), with the incorporation of Principal Component Analysis (PCA) for dimensionality reduction. The ideal quantity of groups (k) for K-means and Agglomerative Filtering is found using the elbow approach. The computational methods are then used, and metric values are used to assess the clustering outcomes. The results indicate that KMeans achieved a Silhouette Score of 0.30, a Calinski-Harabasz Index of 682.37, and a Davies-Bouldin Index of 0.73. GMM yielded a Silhouette Score of 0.29, a Calinski-Harabasz Index of 2.277, and a Davies-Bouldin Index of 365.87. Agglomerative Clustering produced a Silhouette Score of 0.90, a Calinski-Harabasz Index of 331.54, and a Davies-Bouldin Index of 0.87. These results imply that, out of the three categories of algorithms, GMM operated the best. It achieved the highest Silhouette Score and relatively lower Davies-Bouldin Index, indicating better-defined and more separated clusters. All things considered, this work offers insightful information about how methodologies for clustering can be applied to startup data analysis, emphasising the significance of choosing suitable algorithms and assessment metrics in accordance with the features of the dataset and the clustering objectives.