Leveraging Nearest Neighbors for Faithful scHi-C Data Imputation
摘要
Single-cell Hi-C (scHi-C) is a powerful technique for probing the three-dimensional architecture of genomes of individual cells. However, scHi-C data is very sparse, posing computational and analytical challenges for downstream analyses. Existing imputation methods tend to over-impute the data and result in inflated contact frequencies, which may obscure important biological signals. In this paper, we propose a novel imputation method, scHi-C-INN, which leverages the information from the nearest neighbors’ contact matrices. scHi-C-INN preserves the inherent characteristics of the raw data and enhances the quality of scHi-C data. We compare scHi-C-INN with several existing imputation methods on two publicly available scHi-C datasets, and demonstrate its superior performance in terms of cell type clustering and sparseness, and its comparability on TAD-like boundary detection and improved similarities between single cells of the same cell type. The scHi-C-INN will be available at https://github.com/sdontsay/scHi-C-INN .