Sampling Frequent and Diverse Patterns Through Compression
摘要
Exhaustive pattern extraction methods in databases often struggle with speed and output control, leading to the generation of numerous redundant patterns. Sampling-based approaches offer a solution by limiting output size and ensuring faster computation. However, these methods can still produce redundant patterns when a large number is required. For preference learning tasks, it is essential to obtain a concise set of diverse and representative patterns. To address this, we propose integrating compression techniques into the sampling process. This integration refines the selection of representative patterns from sampled transactions while guiding the sampling toward greater diversity. Our approach outperforms existing methods by achieving a more diverse and efficient set of output patterns.