H-QDCT: hierarchical quantum DCT for structural–textural feature fusion in medical imaging
摘要
Deep learning models for medical image classification typically rely on millions of parameters or extensive pretraining, limiting deployment in resource-constrained settings. We investigate whether a hybrid quantum-classical architecture can achieve competitive diagnostic performance with extreme parameter efficiency and superior stability compared to lightweight classical counterparts.
MethodsWe propose the hierarchical quantum discrete cosine transform (H-QDCT), a lightweight pipeline combining classical patch-based preprocessing with quantum frequency-domain feature extraction. Images are partitioned into patches, amplitude-embedded into quantum states, and transformed via a quantum DCT that reorganizes information by spatial frequency. A hierarchical variational ansatz processes structural (low-frequency) and textural (high-frequency) components in disjoint qubits subspaces before global fusion. We evaluate H-QDCT on six clinically binarized MedMNIST datasets under grayscale constraints against both heavy and lightweight classical baselines.
ResultsH-QDCT achieves competitive performance with only 1726 parameters, exceeding a 99.98% reduction compared to ResNet-18 (11.2M parameters). On PneumoniaMNIST, the model attains an AUC of 0.90, approaching the 0.93 of pretrained ResNet-18. Crucially, H-QDCT demonstrates superior robustness in data-scarce regimes: While the lightweight vision transformer (Tiny-ViT) collapsed to the majority class on DermaMNIST (F1
H-QDCT demonstrates that quantum spectral feature extraction achieves clinically meaningful classification with orders-of-magnitude parameter reduction. By mitigating the convergence instability inherent in small-scale classical transformers, H-QDCT establishes a robust design principle for compact, frequency-aware medical diagnostics on near-term quantum hardware.