Shield, Repair, Reconstruct: An Integrated Framework for Long-Term DNA Data Storage
摘要
The exponential growth of global data is rapidly outpacing the capacity of conventional silicon-based storage media, underscoring the need for a new archival solution. DNA, with its unparalleled information density, millennial-scale longevity, and near-zero energy cost for data retention, presents a compelling alternative. However, achieving long-term information storage in DNA hinges on overcoming its central challenge: preserving information integrity against errors introduced during synthesis, storage, amplification, and sequencing. This review synthesizes strategies to maintain data fidelity throughout the DNA storage workflow and argues for a holistic, system-level framework that integrates three layers. First, we examine carrier-level preservation strategies—such as encapsulation in silica, salts, and polymers, as well as in vivo storage—highlighting fundamental trade-offs among stability, density, cost, and random access that create a preservation paradox for system design. Second, we explore biochemical strategies, which actively “repair” accumulated molecular damage prior to sequencing. Third, we describe algorithmic approaches that recover the original data from noisy and fragmented DNA reads through advanced assembly and machine learning methods. We conclude that practical, ultra-long-term DNA storage will emerge from synergistic systems that manage an end-to-end error budget across physical, biochemical, and computational layers. While significant challenges in cost, throughput, and automation remain, the shield-repair-reconstruct paradigm provides a clear roadmap for realizing DNA’s promise as a durable archive of humanity’s digital legacy.