<p>DNA has been proposed as an alternative to magnetic and solid-state devices for storing digital data. In DNA data storage, writing data is performed through DNA synthesis, and reading is done via sequencing. Nanopore devices for sequencing DNA, like those produced by Oxford Nanopore Technologies, allow long reads and real-time sequencing but with lower accuracy compared to other sequencers, such as Illumina. To improve the reliability of data storage in DNA, we aim to combat the high error rate of nanopore sequencing using constrained coding. Certain aspects of the physical process underlying nanopore sequencing mean that some sequences are more prone to sequencing errors than others. We leverage this observation to design constrained codes using constrained de Bruijn graphs, along with a state-splitting encoder and a Viterbi-based decoder. We find that the overall performance of our novel coding system substantially improves upon the state-of-the-art DNN-based methods, reducing sequence-level errors by up to 6 times. We also visually demonstrate the performance of our approach through the simulated recovery of an image encoded and decoded using our method.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Constrained coding for error mitigation in nanopore-based DNA data storage

  • Kallie Whritenour,
  • Mete Civelek,
  • Farzad Farnoud

摘要

DNA has been proposed as an alternative to magnetic and solid-state devices for storing digital data. In DNA data storage, writing data is performed through DNA synthesis, and reading is done via sequencing. Nanopore devices for sequencing DNA, like those produced by Oxford Nanopore Technologies, allow long reads and real-time sequencing but with lower accuracy compared to other sequencers, such as Illumina. To improve the reliability of data storage in DNA, we aim to combat the high error rate of nanopore sequencing using constrained coding. Certain aspects of the physical process underlying nanopore sequencing mean that some sequences are more prone to sequencing errors than others. We leverage this observation to design constrained codes using constrained de Bruijn graphs, along with a state-splitting encoder and a Viterbi-based decoder. We find that the overall performance of our novel coding system substantially improves upon the state-of-the-art DNN-based methods, reducing sequence-level errors by up to 6 times. We also visually demonstrate the performance of our approach through the simulated recovery of an image encoded and decoded using our method.