Temporal Neighboring Multi-modal Transformer with Missingness-Aware Prompt for Hepatocellular Carcinoma Prediction
摘要
Early prediction of hepatocellular carcinoma (HCC) is necessary to facilitate appropriate surveillance strategy and reduce cancer mortality. Incorporating CT scans and clinical time series can greatly increase the accuracy of predictive models. However, there are two challenges to effective multi-modal learning: (a) CT scans and clinical time series suffer from temporal misalignment. (b) CT scans can be missing compared with clinical time series. To tackle the above challenges, we propose a Temporal Neighboring Multi-modal Transformer with Missingness Aware Prompt (TNformer-MP) to integrate clinical time series and available CT scans for HCC prediction. To explore the inter-modality temporal correspondence, a Temporal Neighboring Multi-modal Tokenizer (TN-MT) is exploited to fuse CT embedding into neighboring clinical time series tokens across multiple scales. To mitigate the performance drop caused by missing CT modality, TNformer-MP exploits a Missingness-aware Prompt-driven Multi-modal Tokenizer (MP-MT) that adjusts the encoding of clinical time series tokens with learnable prompts. Experiments conducted on large-scale multi-modal datasets of 36,353 patients show that our method achieves superior performance compared to existing methods.