错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sequential Optimal Experimental Design of Perturbation Screens Guided by Multi-modal Priors

  • Kexin Huang,
  • Romain Lopez,
  • Jan-Christian Hütter,
  • Takamasa Kudo,
  • Antonio Rios,
  • Aviv Regev

摘要

Understanding a cell’s expression response to genetic perturbations helps to address important challenges in biology and medicine. This includes the inference of gene circuits, the discovery of therapeutic targets and the reprogramming and engineering of cells. In recent years, Perturb-seq, pooled genetic screens with single cell RNA-seq (scRNA-seq) readouts, has emerged as a common method to collect such data. Despite technological advancements, the unpredictable, non-additive effects of gene perturbation combinations imply that the number of experimental configurations far exceeds what is experimentally feasible. In some cases, this may even surpass the number of available cells for research. While recent machine learning models, trained on existing Perturb-seq datasets, can predict perturbation outcomes with some degree of accuracy, they are currently limited by sub-optimal training set selection and the small number of cell contexts of training data. This leads to poor predictions for unexplored parts of perturbation space. As biologists deploy Perturb-seq across diverse biological systems, there is a significant need for algorithms to design iterative experiments. These tools are essential for exploring the large space of possible perturbations and their combinations. Here, we propose a sequential approach for designing Perturb-seq experiments that uses the model to strategically select the most informative perturbations at each step for subsequent experiments. This enables a significantly more efficient exploration of the perturbation space, while predicting the effect of the rest of the unseen perturbations with high-fidelity. Analysis of a previous large-scale Perturb-seq experiment reveals that our setting is severely restricted by the number of examples and rounds, falling into a non-conventional active learning regime called “active learning on a budget”. Motivated by this insight, we develop IterPert, a novel active learning method that exploits rich and multi-modal prior knowledge in order to efficiently guide the selection of subsequent perturbations. Using prior knowledge for this task is novel, and crucial for successful active learning on a budget. We validate IterPert using in-silico benchmarking of active learning, constructed from a large-scale CRISPRi Perturb-seq dataset. We find that IterPert outperforms other active learning strategies by reaching comparable accuracy at only a third of the number of perturbations profiled as the next best method. Overall, our results demonstrate the potential of sequentially designing perturbation screens through IterPert.