Homology aware benchmarking of deep learning models for CAZy enzyme family classification
摘要
The CAZy database organizes carbohydrate-active enzymes into more than 800 families with distinct biochemical roles, making automatic family assignment from sequence a high-value task in metagenomics and enzyme engineering. Although recent deep learning studies have reported strong classification performance, we identify a critical evaluation bias that has not been systematically quantified in this setting: commonly used random train/test splits allow homologous sequences to appear in both training and evaluation sets, thereby inflating apparent performance by 5.9–12.0%. We refer to this effect as homology leakage and quantify its magnitude across five representative classification approaches. To enable more rigorous evaluation, we construct a homology-aware benchmark comprising 30,000 sequences, 60 families, and 6 classes, split using MMseqs2 clustering at 20% sequence identity. Under this protocol, the method most strongly affected by leakage is homology-based inference itself (