Ease of access and low cost make the Internet an information hub for many users. But, it comes with many irrelevant contents and web pages, leading to chaos and confusion for the naïve user. The researchers try to develop algorithms to fetch relevant web pages from the web. However, there is scope for improvement of these algorithms by extracting relevant similarity-based features. Existing research utilizes standard similarity measures to find the similarity between relevant and irrelevant textual contents, which leads to less accuracy or bias. Therefore, in this study, authors have proposed a novel page similarity classification algorithm (PSCA) that classifies documents without standard similarity measures. The PSCA is based on content similarity measure (CSM) algorithm, which finds similarities between two web URLs to classify them as true or false. In addition, the authors have generated two novel features, namely new similarity feature (NSF) and set theory-based similarity feature (STBSF) based on algorithms from the literature to test their performance on newly generated datasets. It was observed that all these three algorithms showed higher performance compared to the standard similarity measure algorithm. The proposed PSCA showed an improved performance of 5% compared to all other algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Page Similarity Classification Algorithm for Healthcare Web URL Classification

  • Jatinderkumar R. Saini,
  • Shraddha Vaidya

摘要

Ease of access and low cost make the Internet an information hub for many users. But, it comes with many irrelevant contents and web pages, leading to chaos and confusion for the naïve user. The researchers try to develop algorithms to fetch relevant web pages from the web. However, there is scope for improvement of these algorithms by extracting relevant similarity-based features. Existing research utilizes standard similarity measures to find the similarity between relevant and irrelevant textual contents, which leads to less accuracy or bias. Therefore, in this study, authors have proposed a novel page similarity classification algorithm (PSCA) that classifies documents without standard similarity measures. The PSCA is based on content similarity measure (CSM) algorithm, which finds similarities between two web URLs to classify them as true or false. In addition, the authors have generated two novel features, namely new similarity feature (NSF) and set theory-based similarity feature (STBSF) based on algorithms from the literature to test their performance on newly generated datasets. It was observed that all these three algorithms showed higher performance compared to the standard similarity measure algorithm. The proposed PSCA showed an improved performance of 5% compared to all other algorithms.