Finding the transcription factor binding locations using novel algorithm segmentation to filtration (S2F)
摘要
The primary aim of identifying the binding motifs in gene regulation is to understand the transcriptional regulation molecular mechanism systematically. In this study, the (ℓ, d) motif search issue was considered which entails finding the ℓ length motifs which differ by at most d substitutions. However, identifying the high-quality pattern (ℓ, d) is challenging. It is intended to address the above problem with motif discovery and handle it using the proposed algorithm S2F (Segmentation to Filtration) based on the qPMS (quorum Planted Motif Search) algorithm model. From the entire DNA sequences, five percent are chosen at random to be used in the motif discovery process. This random sub segment (subseg) portion is split up into base, sub k-mers, and its sizes (motif length (ℓ)) are determined by the iterative approach. Corresponding to the sizes of ℓ and d (mutations), the k-mers are chosen which participated in filtration techniques and the base k-mer count and frequency are updated. The highest frequency of k-mer is recognized as the motif. The algorithm’s performance was evaluated using the two real datasets Escherichia coli cyclic AMP receptor protein (CRP) and mouse Embryonic Stem Cell (mESC) ChIP-seq (Chromatin Immuno Precipitation) dataset. Results from the experiments show that S2F can identify the motifs and appear faster compared to previous state-of-the-art PMS (Planted Motif Search) and qPMS algorithms.
Graphical Abstract