<p>In recent years, robust estimation of model parameters has attracted considerable attention in statistics and machine learning, particularly in the context of modeling data with outliers and related inverse problems. This paper introduces a novel log-truncated minimization estimator and a corresponding stochastic gradient descent (SGD) algorithm for quasi-generalized linear models. This approach provides a robust alternative to ordinary GLMs without assuming light-tailed error distributions. For independent non-identical distributed (i.n.i.d.) data, we derive non-asymptotic excess risk and <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_23598_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\ell _2\)</EquationSource> </InlineEquation>-risk bounds using log-truncated Lipschitz losses, assuming only a finite <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_23598_Article_IEq2.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\beta\)</EquationSource> </InlineEquation>-th moment for <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_23598_Article_IEq3.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="68" /> </InlineMediaObject> <EquationSource Format="TEX">\(\beta \in (1,2]\)</EquationSource> </InlineEquation>. Notably, our estimator does not require higher moments or finite variance. Our contribution is to robustify the objective while allowing i.n.i.d. sampling. We also analyze the iteration complexity of SGD for finding stationary points of non-convex log-truncated minimization. Empirically, our SGD algorithm outperforms non-robust methods. We demonstrate its practical effectiveness through a real-data analysis of the German health care demand dataset using robust negative binomial regression.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust learning for ridge-penalized quasi-GLMs under non-identical distributions

  • Huiming Zhang,
  • Wan Tian,
  • Qiuran Yao,
  • Pengfei Wang,
  • Baochang Zhang

摘要

In recent years, robust estimation of model parameters has attracted considerable attention in statistics and machine learning, particularly in the context of modeling data with outliers and related inverse problems. This paper introduces a novel log-truncated minimization estimator and a corresponding stochastic gradient descent (SGD) algorithm for quasi-generalized linear models. This approach provides a robust alternative to ordinary GLMs without assuming light-tailed error distributions. For independent non-identical distributed (i.n.i.d.) data, we derive non-asymptotic excess risk and \(\ell _2\) -risk bounds using log-truncated Lipschitz losses, assuming only a finite \(\beta\) -th moment for \(\beta \in (1,2]\) . Notably, our estimator does not require higher moments or finite variance. Our contribution is to robustify the objective while allowing i.n.i.d. sampling. We also analyze the iteration complexity of SGD for finding stationary points of non-convex log-truncated minimization. Empirically, our SGD algorithm outperforms non-robust methods. We demonstrate its practical effectiveness through a real-data analysis of the German health care demand dataset using robust negative binomial regression.