Navigating the data processing for cytometry-based single-cell proteomics
摘要
Cytometry-based single-cell proteomics (SCP) has emerged as a powerful technique that greatly advances our understanding of complex biological systems with a new level of granularity. Various methods have been developed to process cytometry-based SCP data. However, it remains extremely challenging to identify the well-performing processing workflows for specific datasets. Here, we develop ANPELA, an out-of-the-box method for navigating the proteomic data processing based on large-scale screening. It enables a comparison among the performances of thousands of the processing workflows in identifying cell subpopulations and inferring pseudo-time trajectories based on machine learning. Several cases are then analyzed, highlighting its ability to identify the optimal ways of data processing for cytometry-based SCP studies. A new package is also deployed to ensure multiscenario usability (such as desktop software, R package and online server), data security (enabling local and open-source execution) and a user-friendly interface (realizing interactive and visualizable applications). Overall, ANPELA can be utilized by a broad audience, including those without coding skills, and is freely accessible and downloadable at https://idrblab.org/anpela/. Its execution time may range from minutes to hours depending on the size of the analyzed data.