A knowledge graph-guided transcriptomic framework identifies a 20-gene prognostic signature for cervical cancer
摘要
Current cervical cancer risk stratification relies on FIGO stage and lymph node status, which fail to capture molecular heterogeneity and leave patients over- or under-treated. Conventional transcriptomic biomarker discovery relies solely on differential expression, yielding thousands of candidates that often lack disease specificity. To address these limitations, we developed a computational framework that integrates knowledge graph (KG) embeddings with transcriptomic differential expression to identify cervical cancer-specific prognostic biomarkers.
ResultsAmong ten embedding models for prediction of gene-disease associations, ComplEx showed the highest cervical cancer specificity (Top 100 mean Z-score = 4.19, range 3.54–6.71). Intersection of KG-predicted candidates (Z ≥ 1.645, P < 0.05) with DESeq2 differentially expressed genes (adjusted P < 0.05, |log2FoldChange|> 1) identified 468 high-confidence candidate genes. Using a nested cross-validated Top-K weighted risk score approach, a 20-gene prognostic signature was finalized. The signature demonstrated an unbiased nested cross-validated C-index of 0.638 ± 0.057 for Top-K selection, while the apparent full-data C-index was 0.760 (bootstrap 95% CI 0.697–0.812) and the mean time-dependent AUC was 0.795. Multivariable Cox regression confirmed the signature as an independent prognostic factor after adjusting for age and FIGO stage (HR = 1.89 per SD, P < 0.001, C-index = 0.803), with the strongest discrimination in Stage I disease (C-index = 0.796). Also, external validation in GSE52903 showed a consistent but non-significant trend toward survival discrimination (log-rank P = 0.055; HR = 1.30).
ConclusionsThis study introduces a Z-score-based specificity framework that integrates knowledge-graph embedding with transcriptomic analysis for cervical cancer biomarker discovery. Unlike conventional KG evaluations that rely solely on link-prediction metrics, the disease-specificity Z-score calibrates each gene against ten control solid tumors to filter pan-cancer noise, thereby enhancing the disease specificity of the transcriptomic analysis. These results support the signature’s hypothesis-generating potential for personalized risk stratification, though prospective validation in independent cohorts with survival endpoints remains essential before clinical application.