<p>In response to the shortcomings of current clustering algorithms in achieving high-dimensional text data clustering, a parallel clustering algorithm model for high-dimensional text data was constructed by combining improved K-means and improved s-competitive autoencoder models. First, the s-competitive autoencoder is used to reduce the dimension of high-dimensional text data features, and the extreme learning machine is used to optimize the s-competitive autoencoder model to improve its dimension reduction efficiency and effectiveness. Secondly, a multi-strategy improved dragonfly optimization algorithm is proposed to improve the K-means algorithm. The improved K-means algorithm is called the Pkmeans algorithm, which is used to cluster high-dimensional sparse text data and provide well-preprocessed results for subsequent data processing tasks. Finally, based on the above content, a parallel clustering algorithm model for high-dimensional text data is built. The research outcomes express that compared with the commonly used text feature dimensionality reduction models, the optimized s-competitive autoencoder model has better dimensionality reduction effect. The suggested high-dimensional text data parallel clustering algorithm model outperforms current high-dimensional text clustering algorithms with an accuracy of 98.82% on the BBC dataset, an area under the curve of 94.25%, an Adjusted Rand index of 90.15%, and an NMI of 84.82%. The above findings denote that the proposed method can better achieve parallel clustering of high-dimensional sparse text, which helps to raise the efficiency of text data application.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

High dimensional text data parallel clustering algorithm based on K-means and SAE

  • Jie Zhang

摘要

In response to the shortcomings of current clustering algorithms in achieving high-dimensional text data clustering, a parallel clustering algorithm model for high-dimensional text data was constructed by combining improved K-means and improved s-competitive autoencoder models. First, the s-competitive autoencoder is used to reduce the dimension of high-dimensional text data features, and the extreme learning machine is used to optimize the s-competitive autoencoder model to improve its dimension reduction efficiency and effectiveness. Secondly, a multi-strategy improved dragonfly optimization algorithm is proposed to improve the K-means algorithm. The improved K-means algorithm is called the Pkmeans algorithm, which is used to cluster high-dimensional sparse text data and provide well-preprocessed results for subsequent data processing tasks. Finally, based on the above content, a parallel clustering algorithm model for high-dimensional text data is built. The research outcomes express that compared with the commonly used text feature dimensionality reduction models, the optimized s-competitive autoencoder model has better dimensionality reduction effect. The suggested high-dimensional text data parallel clustering algorithm model outperforms current high-dimensional text clustering algorithms with an accuracy of 98.82% on the BBC dataset, an area under the curve of 94.25%, an Adjusted Rand index of 90.15%, and an NMI of 84.82%. The above findings denote that the proposed method can better achieve parallel clustering of high-dimensional sparse text, which helps to raise the efficiency of text data application.