<p>High-throughput screening data have become the de facto ground truth for in silico toxicology. But a high-throughput “active” can reflect target engagement, non-specific cytotoxicity, or assay interference and the programmes that generate these data flag the latter two with dedicated counter-screens. When models are trained on the hit-calls without those flags, they can learn how the assay behaved rather than how the chemical harms. Because the confounders are structurally systematic, they are exactly the kind of signal a structure-based model will capture. This Commentary sets out the problem and four low-cost controls.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning the assay, not the hazard: how computational toxicity models inherit the artifacts of their high-throughput training data

  • Muhammad Javid Iqbal,
  • Tooba Amjad,
  • Cristian Paz

摘要

High-throughput screening data have become the de facto ground truth for in silico toxicology. But a high-throughput “active” can reflect target engagement, non-specific cytotoxicity, or assay interference and the programmes that generate these data flag the latter two with dedicated counter-screens. When models are trained on the hit-calls without those flags, they can learn how the assay behaved rather than how the chemical harms. Because the confounders are structurally systematic, they are exactly the kind of signal a structure-based model will capture. This Commentary sets out the problem and four low-cost controls.