<p>Bayesian optimization (BO) is a memory-intensive algorithm that requires training and evaluating an expensive objective function. In contrast to previous works that use an offline memory estimation to make BO memory-efficient, we propose a robust and simple online memory estimation method that requires training a model only for the first two iterations of the first epoch. Our memory estimation method is then integrated with a simple, performance-based surrogate model of BO in a seamless (or in sync) mode that enforces memory efficiency even if it does not bypass a preset threshold. The online memory estimation method has been evaluated on two different datasets, showing that it is more accurate than the existing offline method (<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(2.19\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>2.19</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation> for MNIST and <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(3.51\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>3.51</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation> for CIFAR datasets). Furthermore, compared to a memory-unaware baseline, the enhanced BO has no loss of accuracy and is <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(11.31\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>11.31</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation> memory-efficient for a simple CNN-based image classification, and <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(5.03\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>5.03</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation> memory efficient but <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(9.27\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>9.27</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation> slower for a more complex LSTM-based text classification (useful for a resource-constrained environment where delay is tolerable but memory is scarce), while it is <InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(2.6\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>2.6</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation> memory efficient but <InlineEquation ID="IEq7"> <EquationSource Format="TEX">\(1.23\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>1.23</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation> slower on a pretrained ResNet50 model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A memory constrained bayesian optimization via robust online memory estimation

  • Befekadu Bekuretsion,
  • Wolfgang Menzel,
  • Solomon Teferra

摘要

Bayesian optimization (BO) is a memory-intensive algorithm that requires training and evaluating an expensive objective function. In contrast to previous works that use an offline memory estimation to make BO memory-efficient, we propose a robust and simple online memory estimation method that requires training a model only for the first two iterations of the first epoch. Our memory estimation method is then integrated with a simple, performance-based surrogate model of BO in a seamless (or in sync) mode that enforces memory efficiency even if it does not bypass a preset threshold. The online memory estimation method has been evaluated on two different datasets, showing that it is more accurate than the existing offline method ( \(2.19\times \) 2.19 × for MNIST and \(3.51\times \) 3.51 × for CIFAR datasets). Furthermore, compared to a memory-unaware baseline, the enhanced BO has no loss of accuracy and is \(11.31\times \) 11.31 × memory-efficient for a simple CNN-based image classification, and \(5.03\times \) 5.03 × memory efficient but \(9.27\times \) 9.27 × slower for a more complex LSTM-based text classification (useful for a resource-constrained environment where delay is tolerable but memory is scarce), while it is \(2.6\times \) 2.6 × memory efficient but \(1.23\times \) 1.23 × slower on a pretrained ResNet50 model.