<p>It is the purpose of this paper to investigate the issue of estimating the regularity index <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="44" /> </InlineMediaObject> <EquationSource Format="TEX">\(\beta &gt;0\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>β</mi> <mo>&gt;</mo> <mn>0</mn> </mrow> </math></EquationSource> </InlineEquation> of a discrete heavy-tailed r.v. <i>S</i>, <i>i.e.</i> a r.v. <i>S</i> valued in <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(\mathbb {N}^*\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mrow> <mi mathvariant="double-struck">N</mi> </mrow> <mo>∗</mo> </msup> </math></EquationSource> </InlineEquation> such that <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq3.gif" Format="GIF" Height="21" Rendition="HTML" Resolution="72" Type="Linedraw" Width="165" /> </InlineMediaObject> <EquationSource Format="TEX">\(\mathbb {P}(S&gt;n)=L(n)\cdot n^{-\beta }\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi mathvariant="double-struck">P</mi> <mrow> <mo stretchy="false">(</mo> <mi>S</mi> <mo>&gt;</mo> <mi>n</mi> <mo stretchy="false">)</mo> </mrow> <mo>=</mo> <mi>L</mi> <mrow> <mo stretchy="false">(</mo> <mi>n</mi> <mo stretchy="false">)</mo> </mrow> <mo>·</mo> <msup> <mi>n</mi> <mrow> <mo>-</mo> <mi>β</mi> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation> for all <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq4.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(n\ge 1\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>n</mi> <mo>≥</mo> <mn>1</mn> </mrow> </math></EquationSource> </InlineEquation>, where <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq5.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="93" /> </InlineMediaObject> <EquationSource Format="TEX">\(L:\mathbb {R}^*_+\rightarrow \mathbb {R}_+\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>L</mi> <mo>:</mo> <msubsup> <mrow> <mi mathvariant="double-struck">R</mi> </mrow> <mo>+</mo> <mo>∗</mo> </msubsup> <mo stretchy="false">→</mo> <msub> <mi mathvariant="double-struck">R</mi> <mo>+</mo> </msub> </mrow> </math></EquationSource> </InlineEquation> is a slowly varying function. Such discrete probability laws, referred to as generalized Zipf’s laws sometimes, are commonly used to model rank-size distributions after a preliminary range segmentation in a wide variety of areas such as <i>e.g.</i> quantitative linguistics, social sciences or information theory. As a first go, we consider the situation where inference is based on independent copies <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq6.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="84" /> </InlineMediaObject> <EquationSource Format="TEX">\(S_1,\; \ldots ,\; S_n\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msub> <mi>S</mi> <mn>1</mn> </msub> <mo>,</mo> <mspace width="0.277778em" /> <mo>…</mo> <mo>,</mo> <mspace width="0.277778em" /> <msub> <mi>S</mi> <mi>n</mi> </msub> </mrow> </math></EquationSource> </InlineEquation> of the generic variable <i>S</i>. The estimator <InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq7.gif" Format="GIF" Height="22" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\widehat{\beta }\)</EquationSource> <EquationSource Format="MATHML"><math> <mover accent="true"> <mi>β</mi> <mo stretchy="true">^</mo> </mover> </math></EquationSource> </InlineEquation> we propose can be derived by means of a suitable reformulation of the regularly varying condition, replacing <i>S</i>’s survivor function by its empirical counterpart. Under mild assumptions, a non-asymptotic bound for the deviation between <InlineEquation ID="IEq8"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq7.gif" Format="GIF" Height="22" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\widehat{\beta }\)</EquationSource> <EquationSource Format="MATHML"><math> <mover accent="true"> <mi>β</mi> <mo stretchy="true">^</mo> </mover> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq9"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq9.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation> is established, as well as limit results (consistency and asymptotic normality). Beyond the i.i.d. case, the inference method proposed is extended to the estimation of the regularity index of a regenerative <InlineEquation ID="IEq10"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq9.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation>-null-recurrent Markov chain. Since the parameter <InlineEquation ID="IEq11"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11749_2025_975_Article_IEq9.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation> can be then viewed as the tail index of the (regularly varying) distribution of the return time of the chain <i>X</i> to any (pseudo-) regenerative set, in this case, the estimator is constructed from the successive regeneration times. Because the durations between consecutive regeneration times are asymptotically independent, we can prove that the consistency of the estimator promoted is preserved. In addition to the theoretical analysis carried out, simulation results provide empirical evidence of the relevance of the inference technique proposed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tail index estimation for discrete heavy-tailed distributions with application to statistical inference for regular markov chains

  • Patrice Bertail,
  • Stephan Clémençon,
  • Carlos Fernández

摘要

It is the purpose of this paper to investigate the issue of estimating the regularity index \(\beta >0\) β > 0 of a discrete heavy-tailed r.v. S, i.e. a r.v. S valued in \(\mathbb {N}^*\) N such that \(\mathbb {P}(S>n)=L(n)\cdot n^{-\beta }\) P ( S > n ) = L ( n ) · n - β for all \(n\ge 1\) n 1 , where \(L:\mathbb {R}^*_+\rightarrow \mathbb {R}_+\) L : R + R + is a slowly varying function. Such discrete probability laws, referred to as generalized Zipf’s laws sometimes, are commonly used to model rank-size distributions after a preliminary range segmentation in a wide variety of areas such as e.g. quantitative linguistics, social sciences or information theory. As a first go, we consider the situation where inference is based on independent copies \(S_1,\; \ldots ,\; S_n\) S 1 , , S n of the generic variable S. The estimator \(\widehat{\beta }\) β ^ we propose can be derived by means of a suitable reformulation of the regularly varying condition, replacing S’s survivor function by its empirical counterpart. Under mild assumptions, a non-asymptotic bound for the deviation between \(\widehat{\beta }\) β ^ and \(\beta \) β is established, as well as limit results (consistency and asymptotic normality). Beyond the i.i.d. case, the inference method proposed is extended to the estimation of the regularity index of a regenerative \(\beta \) β -null-recurrent Markov chain. Since the parameter \(\beta \) β can be then viewed as the tail index of the (regularly varying) distribution of the return time of the chain X to any (pseudo-) regenerative set, in this case, the estimator is constructed from the successive regeneration times. Because the durations between consecutive regeneration times are asymptotically independent, we can prove that the consistency of the estimator promoted is preserved. In addition to the theoretical analysis carried out, simulation results provide empirical evidence of the relevance of the inference technique proposed.