Image Coding for Storage on Synthetic DNA: Standardization Efforts and Challenges
摘要
DNA-based data storage has recently emerged as a promising paradigm for long-term archival of digital information, offering unparalleled storage density, longevity, and energy efficiency. Among the various classes of digital data, images represent a particularly critical application domain due to their prevalence and high volume, making them well-suited for compression techniques in DNA-based systems. Recognizing these challenges and opportunities, the JPEG Committee initiated the JPEG DNA activity, establishing the first formal standardization effort for the compression of digital images into synthetic DNA sequences.This chapter presents a comprehensive study of the core algorithms and architectural design principles of DNA-based image compression, with a primary emphasis on lossy coding strategies compatible with the biochemical constraints of DNA synthesis, storage, and sequencing. In fact, conventional image compression standards do not natively consider the requirements imposed by the DNA media, highlighting the need for specialized workflows and encoding pipelines to bridge the gap between digital bitstreams and biologically feasible DNA sequences. The chapter focuses in particular on the JPEG DNA Verification Model (VM) pipeline, a codec-agnostic, modular pipeline that includes source coding (adopting JPEG XL as a default coder), a Raptor-based channel encoder with optional error-correction capabilities, and a nucleotide mapping engine that enforces biochemical constraints. Extensive experimental evaluations are conducted to assess the rate—distortion behavior of the JPEG DNA VM under ideal and practical conditions and by adopting multiple source coding methods, including scenarios where oligonucleotides are formatted for wet-lab synthesis and sequencing. Performance comparisons with state-of-the-art methods show how modern codecs such as JPEG XL and JPEG AI offer superior coding efficiency, while legacy codecs like JPEG 2000 provide better tolerance to substitution errors due to their resynchronization capabilities.