Background <p>Single-cell foundation models such as scGPT and Geneformer are increasingly used for gene regulatory network (GRN) inference, with attention-derived edge scores routinely interpreted as regulatory proxies. Prior benchmarks have evaluated curated-reference recovery but have not systematically tested whether attention adds information beyond expression statistics for predicting the outcomes of genetic perturbations, nor whether attention-identified “regulatory” components are causally required for such predictions. This gap matters because the NLP interpretability literature has established that attention weights do not reliably indicate feature importance, and biological foundation models are being deployed without analogous scrutiny.</p> Methods <p>We present an evaluation framework comprising thirty-seven analyses and 153 statistical tests under Benjamini-Hochberg FDR correction, spanning two architectures (scGPT, Geneformer V2-316M), four cell types (K562, RPE1, primary T cells, iPSC neurons), and two perturbation modalities (CRISPRi, CRISPRa). The framework separates two objectives: (A)&#xa0;mechanistic interpretability / GRN recovery against curated references, and (B)&#xa0;perturbation-target prediction, i.e. classifying which genes show differential expression after a CRISPR perturbation. Five test families—trivial-baseline comparison, conditional incremental-value testing, residualisation and propensity matching, causal ablation with intervention-fidelity diagnostics, and cross-context replication—address Objective B, supplemented by a synthetic positive control establishing pipeline sensitivity.</p> Results <p>Attention patterns encode layer-specific biological structure—protein–protein interactions in early layers, transcriptional regulation in late layers—and Cell-State Stratified Interpretability (CSSI) exploits this structure to improve curated GRN recovery up to <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(1.85\times\)</EquationSource> </InlineEquation> on the Objective A task. On Objective B, however, attention-derived edge scores add no incremental value beyond trivial gene-level features (variance, mean expression, dropout rate): gene-level baselines outperform both attention and correlation edges (AUROC 0.81–0.88 versus 0.70), augmenting gene-level predictors with pairwise edges produces <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\Delta\)</EquationSource> </InlineEquation>AUROC of <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(-0.0004\)</EquationSource> </InlineEquation> to <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(-0.002\)</EquationSource> </InlineEquation> across 559,720 perturbation–gene observations, and causal ablation of TRRUST-ranked attention heads produces no degradation across three independent intervention channels. The attention–correlation relationship is context-dependent (equal in K562 CRISPRi, worse in CRISPRa, better in RPE1), but gene-level dominance on Objective B is universal across both cell types where adequate power is available.</p> Conclusions <p>Attention patterns in single-cell foundation models encode biologically structured information, including layer-specific regulatory signals recoverable via CSSI, but provide no unique predictive information beyond simple gene-level statistics for the perturbation-target prediction task. The paper thus offers both a cautionary finding for Objective B and a constructive method (CSSI) for Objective A. Practitioners should apply trivial-baseline and incremental-value tests before claiming pairwise regulatory signal, and should stratify by cell state when extracting attention-derived GRNs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Systematic evaluation of single-cell foundation model interpretability: attention-derived edge scores add no incremental value over gene-level features for perturbation-target prediction

  • Ihor Kendiukhov

摘要

Background

Single-cell foundation models such as scGPT and Geneformer are increasingly used for gene regulatory network (GRN) inference, with attention-derived edge scores routinely interpreted as regulatory proxies. Prior benchmarks have evaluated curated-reference recovery but have not systematically tested whether attention adds information beyond expression statistics for predicting the outcomes of genetic perturbations, nor whether attention-identified “regulatory” components are causally required for such predictions. This gap matters because the NLP interpretability literature has established that attention weights do not reliably indicate feature importance, and biological foundation models are being deployed without analogous scrutiny.

Methods

We present an evaluation framework comprising thirty-seven analyses and 153 statistical tests under Benjamini-Hochberg FDR correction, spanning two architectures (scGPT, Geneformer V2-316M), four cell types (K562, RPE1, primary T cells, iPSC neurons), and two perturbation modalities (CRISPRi, CRISPRa). The framework separates two objectives: (A) mechanistic interpretability / GRN recovery against curated references, and (B) perturbation-target prediction, i.e. classifying which genes show differential expression after a CRISPR perturbation. Five test families—trivial-baseline comparison, conditional incremental-value testing, residualisation and propensity matching, causal ablation with intervention-fidelity diagnostics, and cross-context replication—address Objective B, supplemented by a synthetic positive control establishing pipeline sensitivity.

Results

Attention patterns encode layer-specific biological structure—protein–protein interactions in early layers, transcriptional regulation in late layers—and Cell-State Stratified Interpretability (CSSI) exploits this structure to improve curated GRN recovery up to \(1.85\times\) on the Objective A task. On Objective B, however, attention-derived edge scores add no incremental value beyond trivial gene-level features (variance, mean expression, dropout rate): gene-level baselines outperform both attention and correlation edges (AUROC 0.81–0.88 versus 0.70), augmenting gene-level predictors with pairwise edges produces \(\Delta\) AUROC of \(-0.0004\) to \(-0.002\) across 559,720 perturbation–gene observations, and causal ablation of TRRUST-ranked attention heads produces no degradation across three independent intervention channels. The attention–correlation relationship is context-dependent (equal in K562 CRISPRi, worse in CRISPRa, better in RPE1), but gene-level dominance on Objective B is universal across both cell types where adequate power is available.

Conclusions

Attention patterns in single-cell foundation models encode biologically structured information, including layer-specific regulatory signals recoverable via CSSI, but provide no unique predictive information beyond simple gene-level statistics for the perturbation-target prediction task. The paper thus offers both a cautionary finding for Objective B and a constructive method (CSSI) for Objective A. Practitioners should apply trivial-baseline and incremental-value tests before claiming pairwise regulatory signal, and should stratify by cell state when extracting attention-derived GRNs.