Specificity-driven cell-gene graph learning identifies rare cell states in single-cell and spatial transcriptomic data
摘要
Detecting rare cell populations that drive development, differentiation, and disease-associated transformation remains a central challenge in biology and medicine. Although these populations often represent promising targets for intervention, they are difficult to resolve from single-cell transcriptomic data because most methods rely on homophily-based cell–cell similarity, which can merge rare cells into dominant populations and mask their subtle transcriptional signatures. The challenge is further amplified in multi-sample analyses, where batch correction can dilute rare-cell-specific signals. Here, we present scFormer, a heterogeneous graph transformer (HGT) framework for sensitive and robust rare-cell discovery. scFormer constructs a Z-score-guided cell-gene heterogeneous graph in which highly specific marker genes serve as informational bridges, embedding rare-cell features directly into the graph topology rather than inferring them from global neighbors. This design provides a clear biological rationale for rare-cell recovery, as low-abundance cells can remain connected through shared high-specificity genes even when local cell–cell neighborhoods are sparse. An integrated optimization strategy jointly performs representation learning, clustering, and optional batch correction, enabling rare-cell discovery while preserving biological structure. Across 125 simulated and 18 real datasets, scFormer consistently achieved competitive or superior performance relative to existing approaches. Applied to diverse multi-sample single-cell and spatial transcriptomics datasets, scFormer recovered known but weakly represented populations and revealed previously obscured cell states, including proliferative club cells in the airway epithelium, revival stem cells during intestinal regeneration, and rare embryonic cell states from spatial transcriptomics. Overall, scFormer provides a unified framework for identifying biologically meaningful rare populations while mitigating batch effects in multi-sample datasets.