iHalo: a genome-based machine learning model using k-mer for halophiles identification
摘要
Halophiles hold significant value in both scientific research and industrial applications. However, the identification and experimental validation of halophiles are laborious and time-consuming processes. Consequently, the development of an efficient and convenient method for large-scale halophiles identification is importance. Here, we introduce a novel machine learning approach termed iHalo, capable of directly identifying and mining halophiles from genomic data. To ensure data validity, we constructed a rigorously curated dataset comprising 342 halophilic (with salinity > 15%) and 342 non-halophilic genomes (with salinity < 3%). Subsequently, we conducted a 2–10 k-mer composition analysis, revealing distinct oligonucleotide compositions between halophilic and non-halophilic genomes and identifying three core signature sequences characteristic specific to halophiles. To mitigate noise and enhance computational efficiency, we refined the 87,376 k-mers using an Incremental Feature Selection (IFS) method. Consequently, iHalo achieved its highest accuracy at 175 core features, demonstrating superior predictive accuracy (0.9706) and an area under the curve of 0.9622. Using iHalo, we further mined 1,417 novel halophilic genomes from various genomic databases, with 1,369 of these belonging to Halobacteriales. The remainder were primarily classified as Actinopolysporales, Euryarchaeota, and Halanaerobiales. These findings underscore the efficiency and practicality of iHalo.
Graphical abstract