An Efficient Text-Based Document Categorization with k-Means and Cuckoo Search Optimization
摘要
In the era of information abundance, efficient text-based document retrieval and categorization are of paramount importance. This paper addresses the challenges posed by the exponential growth of digital content and the need for precise and automated document classification. Traditional manual categorization methods are no longer viable due to the enormous data and its structure. Consequently, automated text-based classification techniques, leveraging machine learning, natural language processing, and information retrieval, have gained prominence. Text-based classification faces challenges such as language ambiguity, document diversity, and scalability. A novel approach called k-means and cuckoo search optimization (KCSO) for hierarchical document categorization was proposed in this paper. KCSO combines k-means clustering with optimization techniques to enhance accuracy. Experimental results demonstrate that KCSO outperforms traditional methods, achieving an average accuracy of 90–94%, making it a valuable tool for precise and efficient data categorization.