DNA data storage offers advantages such as high density, and low energy consumption, making it a promising future technology. The incorporation of unnatural nucleic acid bases significantly enhances storage density, stability, and longevity. As a third-generation sequencing technology, nanopore sequencing enables rapid and direct recognition of nucleic acids including both natural and unnatural bases. However, homopolymers in the DNA data storage sequence fail nanopore sequencing due to premature dissociation of DNA molecules from motor protein Hel308 DNA helicase. Here we present a Six-Base Huffman Compression Rotation Coding strategy (S-BHCRC), which uses a six-base system by incorporating two unnatural bases (P, Z) into canonical natural DNA alphabets(A, T, G, C). A rotation coding approach is proposed to eliminate homopolymers in data storage sequences. Meanwhile, we propose a new concept - hydrogen bonds relative content, to replace GC content, as newly introduced unnatural bases (P, Z) pair up with each other with 3 hydrogen bonds. The 6-BHCRC strategy was tested on a data set containing images, videos, and audio files by computer simulations. Results show that average information storage density can reach 2.029 bits/nt, and homopolymers disappear. In summary, by introducing unnatural bases and rotation encoding approach adaptive with nanopore sequencing, 6-BHCRC improves the density of DNA storage and ensures high accuracy in the nanopore sequencing process, demonstrating its immense potential for data storage with ultra-high density and longevity.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

6-Base Huffman Compression Rotation Coding for High Density DNA Data Storage

  • Fei Xu,
  • Sijian Huang,
  • Wanqing Chen

摘要

DNA data storage offers advantages such as high density, and low energy consumption, making it a promising future technology. The incorporation of unnatural nucleic acid bases significantly enhances storage density, stability, and longevity. As a third-generation sequencing technology, nanopore sequencing enables rapid and direct recognition of nucleic acids including both natural and unnatural bases. However, homopolymers in the DNA data storage sequence fail nanopore sequencing due to premature dissociation of DNA molecules from motor protein Hel308 DNA helicase. Here we present a Six-Base Huffman Compression Rotation Coding strategy (S-BHCRC), which uses a six-base system by incorporating two unnatural bases (P, Z) into canonical natural DNA alphabets(A, T, G, C). A rotation coding approach is proposed to eliminate homopolymers in data storage sequences. Meanwhile, we propose a new concept - hydrogen bonds relative content, to replace GC content, as newly introduced unnatural bases (P, Z) pair up with each other with 3 hydrogen bonds. The 6-BHCRC strategy was tested on a data set containing images, videos, and audio files by computer simulations. Results show that average information storage density can reach 2.029 bits/nt, and homopolymers disappear. In summary, by introducing unnatural bases and rotation encoding approach adaptive with nanopore sequencing, 6-BHCRC improves the density of DNA storage and ensures high accuracy in the nanopore sequencing process, demonstrating its immense potential for data storage with ultra-high density and longevity.