Machine Learning and Artificial Intelligence
摘要
Modern AI applications are typically based on machine learning techniques, in particular deep learning which uses multilayered neural networks to simulate the complex decision-making process of the human brain. The evaluation and benchmarking of such techniques is challenging given the broadness of the field and the multitude of possible facets and properties that may be subject to evaluation. In this chapter, we introduce the most popular metrics used in practice to evaluate the quality of machine learning models, covering both classification and regression tasks. We then provide a brief overview of popular machine learning benchmarks and conclude the chapter by briefly discussing the evaluation of self-aware computing systems.