错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Scaling Insights: Unleashing the Power of Scalable Clustering Indices for Big Data Exploration

  • Agnijit Basu,
  • Abhay Narayan

摘要

Data clustering, a fundamental domain in data mining, aims to reveal hidden patterns or groups within datasets. Recent advancements have led to diverse clustering algorithms, whose efficacy often hinges on their computational models. However, as data volumes continue to expand, there is a critical need to validate the outcomes produced by these algorithms. To address this challenge, researchers have explored scalable indexing solutions capable of handling larger datasets efficiently. Our proposed methodology diverges from conventional approaches by employing the “Sketch and Validate” sampling technique. This method aims to approximate scalable index values by leveraging a smaller yet representative data sample. The adoption of scalable models and innovative validation techniques, such as “Sketch and Validate,” plays a pivotal role in addressing the challenges posed by the escalating volume of data in the context of data clustering. To validate our approach, we conducted experiments on two publicly available datasets, with the outcomes meticulously presented in this paper