Modern AI applications are typically based on machine learning techniques, in particular deep learning which uses multilayered neural networks to simulate the complex decision-making process of the human brain. The evaluation and benchmarking of such techniques is challenging given the broadness of the field and the multitude of possible facets and properties that may be subject to evaluation. In this chapter, we introduce the most popular metrics used in practice to evaluate the quality of machine learning models, covering both classification and regression tasks. We then provide a brief overview of popular machine learning benchmarks and conclude the chapter by briefly discussing the evaluation of self-aware computing systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning and Artificial Intelligence

  • André Bauer,
  • Marwin Züfle,
  • Johannes Grohmann,
  • Samuel Kounev

摘要

Modern AI applications are typically based on machine learning techniques, in particular deep learning which uses multilayered neural networks to simulate the complex decision-making process of the human brain. The evaluation and benchmarking of such techniques is challenging given the broadness of the field and the multitude of possible facets and properties that may be subject to evaluation. In this chapter, we introduce the most popular metrics used in practice to evaluate the quality of machine learning models, covering both classification and regression tasks. We then provide a brief overview of popular machine learning benchmarks and conclude the chapter by briefly discussing the evaluation of self-aware computing systems.