<p>Mining implied user similarity is essential for applications such as crowd division and friend recommendation. With the widespread application of cellular signaling data (CSD), it becomes possible to capture users’ daily mobility patterns comprehensively. However, the sparsity and uncertainty of spatiotemporal data pose challenges to accurate similarity measurement. This paper presents a novel historical visit-sequence semantic fusion (HVSSF) framework, which for the first time integrates spatial, semantic, and sequential behavior into a unified structure tailored for CSD analysis. Specifically, HVSSF systematically extracts users’ visit regions, activity types, and behavioral sequences, and then generates semantic embeddings and applies tailored distance metrics to quantify user similarity from both distributional and sequential perspectives. Experiments on three real-world datasets (CSD and GPS) show that our method achieves high accuracy in identifying similar users, with HIT@1 scores of <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41060_2025_817_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(98.7\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>98.7</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41060_2025_817_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(94.2\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>94.2</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, and <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41060_2025_817_Article_IEq3.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(64.0\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>64.0</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, outperforming the best baseline methods by <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41060_2025_817_Article_IEq4.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="35" /> </InlineMediaObject> <EquationSource Format="TEX">\(6.1\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>6.1</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41060_2025_817_Article_IEq5.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(12.6\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>12.6</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, and <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41060_2025_817_Article_IEq6.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="31" /> </InlineMediaObject> <EquationSource Format="TEX">\(14\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>14</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>, respectively. Further evaluations under different observation durations and data completeness levels confirm its robustness. Overall, the HVSSF framework provides a powerful and generalizable solution for user similarity modeling under spatiotemporal uncertainty, offering a new perspective on understanding user behavior from semantic and sequential patterns.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hvssf: A Similar User Mining Method for Cellular Signaling Data With Semantic Information

  • Cheng Zhong,
  • Meng Zhou,
  • Ming Cai

摘要

Mining implied user similarity is essential for applications such as crowd division and friend recommendation. With the widespread application of cellular signaling data (CSD), it becomes possible to capture users’ daily mobility patterns comprehensively. However, the sparsity and uncertainty of spatiotemporal data pose challenges to accurate similarity measurement. This paper presents a novel historical visit-sequence semantic fusion (HVSSF) framework, which for the first time integrates spatial, semantic, and sequential behavior into a unified structure tailored for CSD analysis. Specifically, HVSSF systematically extracts users’ visit regions, activity types, and behavioral sequences, and then generates semantic embeddings and applies tailored distance metrics to quantify user similarity from both distributional and sequential perspectives. Experiments on three real-world datasets (CSD and GPS) show that our method achieves high accuracy in identifying similar users, with HIT@1 scores of \(98.7\%\) 98.7 % , \(94.2\%\) 94.2 % , and \(64.0\%\) 64.0 % , outperforming the best baseline methods by \(6.1\%\) 6.1 % , \(12.6\%\) 12.6 % , and \(14\%\) 14 % , respectively. Further evaluations under different observation durations and data completeness levels confirm its robustness. Overall, the HVSSF framework provides a powerful and generalizable solution for user similarity modeling under spatiotemporal uncertainty, offering a new perspective on understanding user behavior from semantic and sequential patterns.