<p>For DNA data storage, nanopore sequencing can facilitate rapid readout but suffers from severe insertion/deletion errors, which are quite computationally expensive to correct. Here, we propose a nearly single-molecule and assembly-free readout scheme for medium-length pseudo-noise piloting DNA fragments. Specifically, we devise medium-length DNA fragments using low-density parity-check codes companioned by pseudo-noise sequence (PNC-LDPC). A single cleavage on this encoded DNA by transposase generates DNA fragments of approximately full length. Using the readout-aware pseudo-noise sequences, noisy nanopore reads with arbitrary start points are directly located, and base insertions/deletions are corrected, enabling fast and reliable recovery even at very low coverages. Experimental results indicate that the data can be reliably recovered at a coverage of 1.24–3.15× with a typical nanopore sequencing error rate of 1.83%. This method enables error-free recovery in near single-molecule scenarios, highlighting the potential of PNC-LDPC encoded medium-length DNA for data storage applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Approaching single-molecule assembly-free readout from medium-length encoded DNA

  • Weigang Chen,
  • Rui Qin,
  • Quan Guo,
  • Jian Guo,
  • Qi Ge,
  • Yingjin Yuan

摘要

For DNA data storage, nanopore sequencing can facilitate rapid readout but suffers from severe insertion/deletion errors, which are quite computationally expensive to correct. Here, we propose a nearly single-molecule and assembly-free readout scheme for medium-length pseudo-noise piloting DNA fragments. Specifically, we devise medium-length DNA fragments using low-density parity-check codes companioned by pseudo-noise sequence (PNC-LDPC). A single cleavage on this encoded DNA by transposase generates DNA fragments of approximately full length. Using the readout-aware pseudo-noise sequences, noisy nanopore reads with arbitrary start points are directly located, and base insertions/deletions are corrected, enabling fast and reliable recovery even at very low coverages. Experimental results indicate that the data can be reliably recovered at a coverage of 1.24–3.15× with a typical nanopore sequencing error rate of 1.83%. This method enables error-free recovery in near single-molecule scenarios, highlighting the potential of PNC-LDPC encoded medium-length DNA for data storage applications.