<p>Halophiles hold significant value in both scientific research and industrial applications. However, the identification and experimental validation of halophiles are laborious and time-consuming processes. Consequently, the development of an efficient and convenient method for large-scale halophiles identification is importance. Here, we introduce a novel machine learning approach termed iHalo, capable of directly identifying and mining halophiles from genomic data. To ensure data validity, we constructed a rigorously curated dataset comprising 342 halophilic (with salinity &gt; 15%) and 342 non-halophilic genomes (with salinity &lt; 3%). Subsequently, we conducted a 2–10 <i>k</i>-mer composition analysis, revealing distinct oligonucleotide compositions between halophilic and non-halophilic genomes and identifying three core signature sequences characteristic specific to halophiles. To mitigate noise and enhance computational efficiency, we refined the 87,376 <i>k</i>-mers using an Incremental Feature Selection (IFS) method. Consequently, iHalo achieved its highest accuracy at 175 core features, demonstrating superior predictive accuracy (0.9706) and an area under the curve of 0.9622. Using iHalo, we further mined 1,417 novel halophilic genomes from various genomic databases, with 1,369 of these belonging to <i>Halobacteriales</i>. The remainder were primarily classified as <i>Actinopolysporales</i>, <i>Euryarchaeota</i>, and <i>Halanaerobiales</i>. These findings underscore the efficiency and practicality of iHalo.</p> Graphical abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

iHalo: a genome-based machine learning model using k-mer for halophiles identification

  • Li-Hua Liu,
  • Yu Zhang,
  • Wei Lei,
  • Hong Huang,
  • Kuo Zhang,
  • Zhiqian Zhang,
  • Shuqi Wang,
  • Ao Jiang

摘要

Halophiles hold significant value in both scientific research and industrial applications. However, the identification and experimental validation of halophiles are laborious and time-consuming processes. Consequently, the development of an efficient and convenient method for large-scale halophiles identification is importance. Here, we introduce a novel machine learning approach termed iHalo, capable of directly identifying and mining halophiles from genomic data. To ensure data validity, we constructed a rigorously curated dataset comprising 342 halophilic (with salinity > 15%) and 342 non-halophilic genomes (with salinity < 3%). Subsequently, we conducted a 2–10 k-mer composition analysis, revealing distinct oligonucleotide compositions between halophilic and non-halophilic genomes and identifying three core signature sequences characteristic specific to halophiles. To mitigate noise and enhance computational efficiency, we refined the 87,376 k-mers using an Incremental Feature Selection (IFS) method. Consequently, iHalo achieved its highest accuracy at 175 core features, demonstrating superior predictive accuracy (0.9706) and an area under the curve of 0.9622. Using iHalo, we further mined 1,417 novel halophilic genomes from various genomic databases, with 1,369 of these belonging to Halobacteriales. The remainder were primarily classified as Actinopolysporales, Euryarchaeota, and Halanaerobiales. These findings underscore the efficiency and practicality of iHalo.

Graphical abstract