Approximate Density Computation for OA-Biclustering
摘要
In this paper, we introduce an approach for approximate density computation of object-attribute biclusters (OA-biclusters) given in the form \((m',g')\) , where \((.)'\) stands for Galois operators from Formal Concept Analysis. They serve as an approximate variant of formal concepts since not all object-attribute pairs for the original incidence relation are present in \(m'\times g'\) . Direct density computation of an OA-bicluster takes \(O(|g'||m'|)\) , which is laborious for large biclusters, while its high density serves as a high-quality guarantee (in the sense of interestingness measure). Our approach relies on the Monte Carlo sampling strategy and concentration inequalities known in statistics and machine learning. It is tested on different datasets including Southern Women from Social Networks Analysis domain, the Zoo dataset from the UCI ML repository, custom artificial genotype-phenotype data from genome association study, and web advertising data donated by Yahoo (formerly US Overture). The results of the experimental comparison show a reasonably good quality of approximation in terms of average error and demonstrate the possible applications of the approach in different domains.