Bioinformatics on Sperm Subpopulations Using Computer-Assisted Sperm Analysis (CASA)
摘要
Computer-aided sperm analysis of motility (CASA-mot) and morphology (CASA-morph) enables the quick acquisition of many parameters on individual spermatozoa. This results in massive datasets apt for exploration using multiparametric tools. Data clustering has been used for decades to identify subpopulations of spermatozoa sharing similar motility patterns or morphological features. There are two main strategies considering the lack of or availability of prior information: unsupervised and supervised clustering. Then, the researcher has a wealth of algorithms for processing the data in both approaches, but one has to be aware that not all of them could be appropriate for dealing with the peculiarities of sperm data. Here, we describe in detail two approaches previously published: the two-step procedure as unsupervised and the support vector machines (SVM) as supervised. Both approaches are widely available in different statistical software, but in the interest of openness in research, specific details are provided for the R statistical environment (open source and non-proprietary platform and packages). This chapter offers strengths, caveats, and suggestions for alternative algorithms.