Background <p>Variational autoencoders (VAEs) serve as essential components in large generative models for extracting latent representations and have gained widespread application in biological domains. Developing VAEs specifically tailored to the unique characteristics of biological data is crucial for advancing future large-scale biological models.</p> Results <p>Through systematic monitoring of VAE training processes across 31 public single-cell datasets spanning oncological and normal conditions, we discovered that reducing the <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="12915_2025_2315_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\beta\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation> value which corresponds to lower disentanglement of VAE significantly improves unsupervised clustering metrics in single-cell data analysis. Based on this finding, we innovatively developed iVAE with an irecon module that, when benchmarked against 8 established dimensionality reduction methods across 5 clustering performance metrics, exhibited superior capabilities in representing single-cell transcriptomic data.</p> Conclusions <p>The proposed iVAE architecture enhances the interpretability of single-cell data compared to conventional VAE architectures as measured by clustering metrics. Our work establishes a potential foundational VAE architecture for developing specialized large-scale generative models for biological applications.</p> Graphical abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

iVAE: an interpretable representation learning framework enhances clustering performance for single-cell data

  • Zeyu Fu,
  • Chunlin Chen,
  • Song Wang,
  • Junping Wang,
  • Shilei Chen

摘要

Background

Variational autoencoders (VAEs) serve as essential components in large generative models for extracting latent representations and have gained widespread application in biological domains. Developing VAEs specifically tailored to the unique characteristics of biological data is crucial for advancing future large-scale biological models.

Results

Through systematic monitoring of VAE training processes across 31 public single-cell datasets spanning oncological and normal conditions, we discovered that reducing the \(\beta\) β value which corresponds to lower disentanglement of VAE significantly improves unsupervised clustering metrics in single-cell data analysis. Based on this finding, we innovatively developed iVAE with an irecon module that, when benchmarked against 8 established dimensionality reduction methods across 5 clustering performance metrics, exhibited superior capabilities in representing single-cell transcriptomic data.

Conclusions

The proposed iVAE architecture enhances the interpretability of single-cell data compared to conventional VAE architectures as measured by clustering metrics. Our work establishes a potential foundational VAE architecture for developing specialized large-scale generative models for biological applications.

Graphical abstract