T Cell Receptor Protein Sequences and Sparse Coding: A Novel Approach to Cancer Classification
摘要
Cancer is a complex disease marked by uncontrolled cell growth, potentially leading to tumors and metastases. Identifying cancer types is crucial for treatment decisions and patient outcomes. T Cell receptors (TCRs) are vital proteins in adaptive immunity, specifically recognizing antigens and playing a pivotal role in immune responses, including against cancer. TCR diversity makes them promising for targeting cancer cells, aided by advanced sequencing revealing potent anti-cancer TCRs and TCR-based therapies. Effectively analyzing these complex biomolecules necessitates representation and capturing their structural and functional essence. We explore sparse coding for multi-classifying TCR protein sequences with cancer categories as targets. Sparse coding, a machine learning technique, represents data with informative features, capturing intricate amino acid relationships and subtle sequence patterns. We compute TCR sequence k-mers, applying sparse coding to extract key features. Domain knowledge integration improves predictive embeddings, incorporating cancer properties like Human leukocyte antigen (HLA) types, gene mutations, clinical traits, immunological features, and epigenetic changes. Our embedding method, applied to a TCR benchmark dataset, significantly outperforms baselines, achieving 99.8% accuracy. Our study underscores sparse coding’s potential in dissecting TCR protein sequences in cancer research.