Cluster analysis organizes cells into clusters based on the similarity of single-cell gene expression profiles, representing unique cell states or types in the transcriptome space, in order to infer and study cell heterogeneity, define and discover cell types. Faced with the high complexity, high-dimensional sparsity, and high noise of high-throughput scRNA-seq data caused by their respective biological and technological impacts, most existing methods directly use processed single-cell data for neural network training, without mining the feature information of the original data. In this study, we have proposed a deep learning clustering framework (scLRGCN) based on low rank representation for mining overlooked and confused single-cell data structures and information. With various low rank representation algorithms achieving good results in subspace clustering, we introduce low rank representation for similarity learning. The similarity matrix learned from single-cell raw data is used for training Graph Convolutional Networks (GCNs) to better address high-dimensional sparsity issues through a raw data-driven approach. We also added a denoising single-cell graph data construction method for metadata to enhance the shielding of noise effects Finally, each module is supervised and trained through target distribution to optimize clustering performance. Experimental results have shown that the model has achieved good results on various datasets and clustering evaluation indicators, with certain generalization ability and robustness, which plays an important role and significance in improving clustering performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Deep Clustering Method for Single-Cell Data Based on Low Rank Representation

  • Zeran You,
  • Bo Li

摘要

Cluster analysis organizes cells into clusters based on the similarity of single-cell gene expression profiles, representing unique cell states or types in the transcriptome space, in order to infer and study cell heterogeneity, define and discover cell types. Faced with the high complexity, high-dimensional sparsity, and high noise of high-throughput scRNA-seq data caused by their respective biological and technological impacts, most existing methods directly use processed single-cell data for neural network training, without mining the feature information of the original data. In this study, we have proposed a deep learning clustering framework (scLRGCN) based on low rank representation for mining overlooked and confused single-cell data structures and information. With various low rank representation algorithms achieving good results in subspace clustering, we introduce low rank representation for similarity learning. The similarity matrix learned from single-cell raw data is used for training Graph Convolutional Networks (GCNs) to better address high-dimensional sparsity issues through a raw data-driven approach. We also added a denoising single-cell graph data construction method for metadata to enhance the shielding of noise effects Finally, each module is supervised and trained through target distribution to optimize clustering performance. Experimental results have shown that the model has achieved good results on various datasets and clustering evaluation indicators, with certain generalization ability and robustness, which plays an important role and significance in improving clustering performance.