In this work we introduce DATE (Derivative Alignment Training for Extrapolation), a method to improve the extrapolation behaviour of neural networks (NN) with Rectified Linear Unit activation (ReLU) on univariate regression tasks. ReLU NNs naturally lend themselves to linear extrapolation beyond the training data range. However, there are two known limitations of extrapolation properties of trained ReLU NNs, that we address in this paper. When minimising the error of the prediction, the derivative of the NN model function can still show high variation, which can cause variable extrapolation. Non-linearities of the model function outside the training data range can lead to inconsistent extrapolation behaviour. In prior work, the extrapolation issue has been addressed with a set of regularisation functions, called ReLEx. To improve extrapolation and interpolation, we introduce two new regularisation terms: D1-loss and IR-loss. The D1-loss directly penalises the deviation of the model derivative from a target derivative as estimated from the data by interpolating between neighbouring data points. The IR-loss penalises positions of the non-linearities of the ReLU units outside a given range. Optimising the combination of D1 with IR loss and/or some of the ReLEx functions constitutes the DATE method. We evaluate DATE on regression tasks with noiseless data generated from analytic functions. We test different DATE configurations and find that training with DATE can reduce the variability of the model slope, prevent non-linearities outside the training data range, and improve extrapolation consistency as measured by different metrics. The most effective DATE variants also have reduced complexity compared to ReLEx.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DATE: Derivative Alignment Training for Extrapolation with Neural Networks

  • Enrico Lopedoto,
  • Tillman Weyde,
  • Kizito Salako

摘要

In this work we introduce DATE (Derivative Alignment Training for Extrapolation), a method to improve the extrapolation behaviour of neural networks (NN) with Rectified Linear Unit activation (ReLU) on univariate regression tasks. ReLU NNs naturally lend themselves to linear extrapolation beyond the training data range. However, there are two known limitations of extrapolation properties of trained ReLU NNs, that we address in this paper. When minimising the error of the prediction, the derivative of the NN model function can still show high variation, which can cause variable extrapolation. Non-linearities of the model function outside the training data range can lead to inconsistent extrapolation behaviour. In prior work, the extrapolation issue has been addressed with a set of regularisation functions, called ReLEx. To improve extrapolation and interpolation, we introduce two new regularisation terms: D1-loss and IR-loss. The D1-loss directly penalises the deviation of the model derivative from a target derivative as estimated from the data by interpolating between neighbouring data points. The IR-loss penalises positions of the non-linearities of the ReLU units outside a given range. Optimising the combination of D1 with IR loss and/or some of the ReLEx functions constitutes the DATE method. We evaluate DATE on regression tasks with noiseless data generated from analytic functions. We test different DATE configurations and find that training with DATE can reduce the variability of the model slope, prevent non-linearities outside the training data range, and improve extrapolation consistency as measured by different metrics. The most effective DATE variants also have reduced complexity compared to ReLEx.