错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pan-cell type continuous chromatin state annotation of all epigenomes from the International Human Epigenome Consortium

  • Habib Daneshpajouh,
  • Ismail Moghul,
  • Kay C. Wiese,
  • Maxwell W. Libbrecht

摘要

Background

The International Human Epigenome Consortium has generated thousands of datasets profiling transcription factor binding, histone modification, and DNA accessibility. Existing segmentation and annotation approaches typically produce cell type-specific regulatory maps, which have become increasingly difficult to maintain and apply as the number of profiled cell types has grown. Here, we apply epigenome-ssm, a continuous state-space modeling framework, to generate a unified, interpretable representation of chromatin states across thousands of human epigenomes.

Results

Using 9,539 histone modification signal tracks from 1,698 epigenomes, epigenome-ssm produces 33 continuous chromatin state features that compactly capture regulatory programs across cell types. These features distinguish canonical activities such as promoters, enhancers, transcription, and heterochromatin, while also encoding cell type-specific regulatory patterns. Compared with alternative pan-cell type annotation methods, the continuous features achieve superior or comparable predictive performance for gene expression, enhancer activity, and evolutionary conservation, despite using fewer dimensions. The model effectively captures both broad and lineage-specific regulatory programs, linking chromatin states to gene expression and functional annotations. Additionally, a derived conservation-associated activity score (SSM-CAAS) highlights genomic regions enriched for disease-associated variants, demonstrating utility for interpreting noncoding variation.

Conclusions

Continuous pan-cell type chromatin state features provide a compact, expressive, and biologically informative representation of the human epigenome. This framework improves integration and interpretation of large-scale epigenomic data, enables accurate prediction of genomic function, and facilitates identification of regulatory elements relevant to disease. The resulting resource offers a scalable foundation for downstream analyses of gene regulation and genetic variation.