<p>Single-cell RNA sequencing (scRNA-seq) technology enables a deep understanding of cellular differentiation during plant development and reveals heterogeneity among the cells of a given tissue. However, the computational characterization of such cellular heterogeneity is complicated by the high dimensionality, sparsity, and biological noise inherent to the raw data. Here, we introduce PhytoCluster, an unsupervised deep learning algorithm, to cluster scRNA-seq data by extracting latent features. We benchmarked PhytoCluster against four simulated datasets and five real scRNA-seq datasets with varying protocols and data quality levels. A comprehensive evaluation indicated that PhytoCluster outperforms other methods in clustering accuracy, noise removal, and signal retention. Additionally, we evaluated the performance of the latent features extracted by PhytoCluster across four machine learning models. The computational results highlight the ability of PhytoCluster to extract meaningful information from plant scRNA-seq data, with machine learning models achieving accuracy comparable to that of raw features. We believe that PhytoCluster will be a valuable tool for disentangling complex cellular heterogeneity based on scRNA-seq data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PhytoCluster: a generative deep learning model for clustering plant single-cell RNA-seq data

  • Hao Wang,
  • Xiangzheng Fu,
  • Lijia Liu,
  • Yi Wang,
  • Jingpeng Hong,
  • Bintao Pan,
  • Yaning Cao,
  • Yanqing Chen,
  • Yongsheng Cao,
  • Xiaoding Ma,
  • Wei Fang,
  • Shen Yan

摘要

Single-cell RNA sequencing (scRNA-seq) technology enables a deep understanding of cellular differentiation during plant development and reveals heterogeneity among the cells of a given tissue. However, the computational characterization of such cellular heterogeneity is complicated by the high dimensionality, sparsity, and biological noise inherent to the raw data. Here, we introduce PhytoCluster, an unsupervised deep learning algorithm, to cluster scRNA-seq data by extracting latent features. We benchmarked PhytoCluster against four simulated datasets and five real scRNA-seq datasets with varying protocols and data quality levels. A comprehensive evaluation indicated that PhytoCluster outperforms other methods in clustering accuracy, noise removal, and signal retention. Additionally, we evaluated the performance of the latent features extracted by PhytoCluster across four machine learning models. The computational results highlight the ability of PhytoCluster to extract meaningful information from plant scRNA-seq data, with machine learning models achieving accuracy comparable to that of raw features. We believe that PhytoCluster will be a valuable tool for disentangling complex cellular heterogeneity based on scRNA-seq data.