错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On the Minimum Error Using Kolmogorov Size Shallow Neural Network and Gradient Descent Algorithms for Complicated Univariate Functions

  • J. Pablo Altamirano,
  • Alejandro Santiago,
  • José Antonio Castán Rocha,
  • Julisa Pérez Cobos,
  • Roberto Pichardo Ramírez

摘要

Artificial Neural Networks (ANNs) are among the most influential research topics in Artificial Intelligence (AI). ANNs are the fundamental base for recent advances in Generative AI for text, image, music, code, and voice generation. Several research studies focus on Deep Neural Networks (DNNs) to perform the most challenging tasks, such as natural language processing, i.e., large language models. Nevertheless, DNNs have increased the computer power required to operate them, having even billions of parameters. Although Shallow Artificial Neural Networks have no theoretical limits to perform as universal approximations, current research is leadership by DNN instead of Shallow ANNs. This chapter focuses on experimentally evaluating Shallow Neural Networks with \(2n+1\) hidden units, which, according to Kolmogorov’s theorem, are the necessary number of units to implement any function. We selected a set of 18 univariate complicated shape functions and 13 gradient descent-based algorithms: Batch Gradient Descent, Stochastic Gradient Descent, Minibatch Gradient Descent, Momentum Stochastic Gradient Descent, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, MAXprop, Adam, AdaMax, Nadam, NadaMax for experimental purposes. Given their lower requirement for computer power than DNN and no theoretical limits, research on Shallow Neural Networks is of great interest. Experiments show that gradient descent-based methods are far from optimality.