Learning the assay, not the hazard: how computational toxicity models inherit the artifacts of their high-throughput training data
摘要
High-throughput screening data have become the de facto ground truth for in silico toxicology. But a high-throughput “active” can reflect target engagement, non-specific cytotoxicity, or assay interference and the programmes that generate these data flag the latter two with dedicated counter-screens. When models are trained on the hit-calls without those flags, they can learn how the assay behaved rather than how the chemical harms. Because the confounders are structurally systematic, they are exactly the kind of signal a structure-based model will capture. This Commentary sets out the problem and four low-cost controls.