Background <p>DNA replication is a biological process in which a single DNA molecule is duplicated, initiating from multiple genomic sites known as replication origins. Identifying replication origins and analyzing their underlying base sequence composition is crucial for understanding the mechanisms of DNA replication. Although there are various machine learning and deep learning approaches for origin prediction, many rely on labor intensive feature engineering or lack interpretability. We fine-tune two genome-based pretrained language models, DNABERT and DNABERT-2, to predict replication origins in budding yeast and unravel the DNA base composition behind them. The key contribution of this study is a systematic framework for analyzing genomic language models for replication origin prediction, combining controlled dataset design with model-specific explainability pipelines to examine how different tokenization strategies influence learned sequence features and whether such approaches can highlight biologically meaningful signals.</p> Results <p>We evaluate both models on the designed datasets to ensure robustness and support explainability. DNABERT demonstrates consistent performance, achieving an average accuracy of 0.72 for more challenging and 0.83 for the easier dataset. In comparison, DNABERT-2 achieved comparable scores of 0.72 and 0.81 on the same datasets. Our attention-based motif discovery pipeline enhances the interpretability of DNABERT, by identifying motifs from high-attention fragments that closely match known sequence patterns of replication origins. Perturbation-based explanation methods, including Shapley additive explanations, were applied to interpret DNABERT-2’s learning mechanism. This analysis identified tokens with high attribution scores aligned with biologically relevant sequence composition.</p> Conclusion <p>Our study demonstrates that both models identify replication origin sequences, albeit through different learning strategies. Tokenization appears to influence model learning and attention behavior in these models. The overlapping k-mer tokenization used in DNABERT yields more interpretable attention maps compared to the byte pair encoding tokenization employed in DNABERT-2. We show that despite sharing the same BERT-style architecture, DNABERT captures relevant short-range patterns and some sequence dependencies beyond just local context, as reflected in its attention maps. In contrast, DNABERT-2’s alternative tokenization strategy biases its learning toward relevant short-range patterns by optimizing token weighting.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Interpretable prediction of DNA replication origins in S. cerevisiae using DNABERT and DNABERT-2

  • Zohreh Piroozeh,
  • Ildem Akerman,
  • Olga V. Kalinina,
  • Stefan Kesselheim,
  • Alina Bazarova

摘要

Background

DNA replication is a biological process in which a single DNA molecule is duplicated, initiating from multiple genomic sites known as replication origins. Identifying replication origins and analyzing their underlying base sequence composition is crucial for understanding the mechanisms of DNA replication. Although there are various machine learning and deep learning approaches for origin prediction, many rely on labor intensive feature engineering or lack interpretability. We fine-tune two genome-based pretrained language models, DNABERT and DNABERT-2, to predict replication origins in budding yeast and unravel the DNA base composition behind them. The key contribution of this study is a systematic framework for analyzing genomic language models for replication origin prediction, combining controlled dataset design with model-specific explainability pipelines to examine how different tokenization strategies influence learned sequence features and whether such approaches can highlight biologically meaningful signals.

Results

We evaluate both models on the designed datasets to ensure robustness and support explainability. DNABERT demonstrates consistent performance, achieving an average accuracy of 0.72 for more challenging and 0.83 for the easier dataset. In comparison, DNABERT-2 achieved comparable scores of 0.72 and 0.81 on the same datasets. Our attention-based motif discovery pipeline enhances the interpretability of DNABERT, by identifying motifs from high-attention fragments that closely match known sequence patterns of replication origins. Perturbation-based explanation methods, including Shapley additive explanations, were applied to interpret DNABERT-2’s learning mechanism. This analysis identified tokens with high attribution scores aligned with biologically relevant sequence composition.

Conclusion

Our study demonstrates that both models identify replication origin sequences, albeit through different learning strategies. Tokenization appears to influence model learning and attention behavior in these models. The overlapping k-mer tokenization used in DNABERT yields more interpretable attention maps compared to the byte pair encoding tokenization employed in DNABERT-2. We show that despite sharing the same BERT-style architecture, DNABERT captures relevant short-range patterns and some sequence dependencies beyond just local context, as reflected in its attention maps. In contrast, DNABERT-2’s alternative tokenization strategy biases its learning toward relevant short-range patterns by optimizing token weighting.