<p>Rapid prediction of antimicrobial resistance (AMR) from genome sequences is essential for timely therapy, yet models based on curated marker panels or core-genome Single Nucleotide Polymorphisms (SNPs) often fail to generalize to novel bacterial lineages. We evaluate AMR prediction in <i>Staphylococcus aureus</i> using pan-genome features that encode homologous gene copy number (including absence) and compare them to SNP-based models across six antibiotics and 4255 isolates. Gradient-boosted decision tree ensembles (XGBoost) trained on gene copy number achieve macro-averaged F1-scores of 0.925–0.988, surpassing SNP-based models (0.838–0.935). Under lineage-held-out evaluation, which withholds entire clades to mimic previously unseen lineages, gene-content models retain markedly higher performance (F1 = 0.875 and 0.904 across two split schemes), whereas SNP-based models degrade substantially (F1 = 0.557 and 0.638). Feature ablation indicates that predictive signal is distributed across many homologous gene families rather than dominated by a few markers, a structure consistent with stronger cross-lineage generalization. Because gene-content features can be robustly obtained even from low-coverage sequencing, this approach extends genome-based AMR prediction to real-world clinical and epidemiological datasets. Together, these results show that copy-number-based pan-genome representations provide a robust alternative to SNP-only approaches, particularly when models must generalize to lineages not represented in training data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Gene copy-number features generalize better than SNPs for antimicrobial resistance prediction in Staphylococcus aureus

  • Bruna F. Fistarol,
  • Joao D. Gervasio,
  • Gergely J. Szöllősi

摘要

Rapid prediction of antimicrobial resistance (AMR) from genome sequences is essential for timely therapy, yet models based on curated marker panels or core-genome Single Nucleotide Polymorphisms (SNPs) often fail to generalize to novel bacterial lineages. We evaluate AMR prediction in Staphylococcus aureus using pan-genome features that encode homologous gene copy number (including absence) and compare them to SNP-based models across six antibiotics and 4255 isolates. Gradient-boosted decision tree ensembles (XGBoost) trained on gene copy number achieve macro-averaged F1-scores of 0.925–0.988, surpassing SNP-based models (0.838–0.935). Under lineage-held-out evaluation, which withholds entire clades to mimic previously unseen lineages, gene-content models retain markedly higher performance (F1 = 0.875 and 0.904 across two split schemes), whereas SNP-based models degrade substantially (F1 = 0.557 and 0.638). Feature ablation indicates that predictive signal is distributed across many homologous gene families rather than dominated by a few markers, a structure consistent with stronger cross-lineage generalization. Because gene-content features can be robustly obtained even from low-coverage sequencing, this approach extends genome-based AMR prediction to real-world clinical and epidemiological datasets. Together, these results show that copy-number-based pan-genome representations provide a robust alternative to SNP-only approaches, particularly when models must generalize to lineages not represented in training data.