<p>In Bayesian inference, the asymptotic behavior of evaluation metrics for predictive accuracy, such as generalization loss and free energy, is determined by a model-specific rational number known as the learning coefficient. Learning coefficients are already known for models that satisfy regularity conditions. However, for singular models that do not satisfy these conditions, specific values of learning coefficients have been provided for certain models such as reduced-rank regression. Nevertheless, a general formula for learning coefficients that broadly applies to singular models has not yet been established. Kurumadani proposed a formula for the learning coefficient of singular models called semiregular models. However, it can only be used when the Kullback–Leibler divergence <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(K(\theta )\)</EquationSource> </InlineEquation> is analytic or when the set of realization parameters <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\Theta _*\)</EquationSource> </InlineEquation> is a single point. It is not applicable to more general cases. In this paper, we present two generalizations to overcome this limitation. First, we show that the formula for the learning coefficient still holds when <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(K(\theta )\)</EquationSource> </InlineEquation> is not analytic. Second, we derive a formula that is applicable even when <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\Theta _*\)</EquationSource> </InlineEquation> is not a single point. As a result, we gain a systematic understanding of how the natural number <i>m</i>, determined by partial derivatives of the log-likelihood ratio function, and the dimension of <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(\Theta _*\)</EquationSource> </InlineEquation>, a geometric quantity, influence the learning coefficient. Our main results are applicable to a wide range of singular models that were not addressed in Kurumadani. As a specific example, we provide upper bounds on the learning coefficients for neural networks and mixture distributions, which had not been elucidated before.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Learning coefficients in semiregular models II: extensions

  • Yuki Kurumadani

摘要

In Bayesian inference, the asymptotic behavior of evaluation metrics for predictive accuracy, such as generalization loss and free energy, is determined by a model-specific rational number known as the learning coefficient. Learning coefficients are already known for models that satisfy regularity conditions. However, for singular models that do not satisfy these conditions, specific values of learning coefficients have been provided for certain models such as reduced-rank regression. Nevertheless, a general formula for learning coefficients that broadly applies to singular models has not yet been established. Kurumadani proposed a formula for the learning coefficient of singular models called semiregular models. However, it can only be used when the Kullback–Leibler divergence \(K(\theta )\) is analytic or when the set of realization parameters \(\Theta _*\) is a single point. It is not applicable to more general cases. In this paper, we present two generalizations to overcome this limitation. First, we show that the formula for the learning coefficient still holds when \(K(\theta )\) is not analytic. Second, we derive a formula that is applicable even when \(\Theta _*\) is not a single point. As a result, we gain a systematic understanding of how the natural number m, determined by partial derivatives of the log-likelihood ratio function, and the dimension of \(\Theta _*\) , a geometric quantity, influence the learning coefficient. Our main results are applicable to a wide range of singular models that were not addressed in Kurumadani. As a specific example, we provide upper bounds on the learning coefficients for neural networks and mixture distributions, which had not been elucidated before.