Deep learning as a bridge between sequencing-derived biological data and functional interpretation
摘要
High-throughput sequencing has generated large-scale genomic, transcriptomic, and multimodal datasets, but converting these data into biological interpretation remains difficult because of dimensionality, sparsity, noise, and incomplete ground truth. Deep learning offers flexible representation-learning frameworks for modeling local sequence motifs, long-range dependencies, relational structures, and latent cellular states. In this narrative review, we summarize convolutional, recurrent, graph-based, generative, and transformer architectures and examine representative applications in regulatory genomics, RNA structure prediction, single-cell and spatial transcriptomics, pathology-linked multimodal inference, and proteome-related prediction tasks. We emphasize that these models mainly produce statistical predictions, learned representations, and candidate regulatory signals; they can support functional hypotheses but do not establish biological function without independent validation. We also discuss recurring limitations, including dataset and annotation bias, domain shift, tokenization choices, computational cost, limited reproducibility, incomplete uncertainty quantification, and weak causal identifiability. Finally, we discuss artificial-intelligence virtual cells as an emerging conceptual objective rather than an established capability. Overall, deep learning is framed as a powerful tool for pattern discovery and hypothesis generation, whose biological interpretation requires careful benchmarking, external validation, and experimental follow-up.