Purpose <p>Deep learning models for medical image classification typically rely on millions of parameters or extensive pretraining, limiting deployment in resource-constrained settings. We investigate whether a hybrid quantum-classical architecture can achieve competitive diagnostic performance with extreme parameter efficiency and superior stability compared to lightweight classical counterparts.</p> Methods <p>We propose the hierarchical quantum discrete cosine transform (H-QDCT), a lightweight pipeline combining classical patch-based preprocessing with quantum frequency-domain feature extraction. Images are partitioned into patches, amplitude-embedded into quantum states, and transformed via a quantum DCT that reorganizes information by spatial frequency. A hierarchical variational ansatz processes structural (low-frequency) and textural (high-frequency) components in disjoint qubits subspaces before global fusion. We evaluate H-QDCT on six clinically binarized MedMNIST datasets under grayscale constraints against both heavy and lightweight classical baselines.</p> Results <p>H-QDCT achieves competitive performance with only 1726 parameters, exceeding a 99.98% reduction compared to ResNet-18 (11.2M parameters). On PneumoniaMNIST, the model attains an AUC of 0.90, approaching the 0.93 of pretrained ResNet-18. Crucially, H-QDCT demonstrates superior robustness in data-scarce regimes: While the lightweight vision transformer (Tiny-ViT) collapsed to the majority class on DermaMNIST (F1 <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(=0.00\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>=</mo> <mn>0.00</mn> </mrow> </math></EquationSource> </InlineEquation>) and BloodMNIST (F1 <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(=0.06\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>=</mo> <mn>0.06</mn> </mrow> </math></EquationSource> </InlineEquation>), H-QDCT retained learnability on BloodMNIST (F1 up to 0.22) with stable convergence. Beyond binary screening, the architecture extends natively to multi-class differential diagnosis (from <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(K=2\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>K</mi> <mo>=</mo> <mn>2</mn> </mrow> </math></EquationSource> </InlineEquation> to <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(K=9\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>K</mi> <mo>=</mo> <mn>9</mn> </mrow> </math></EquationSource> </InlineEquation>) without mode collapse, and on PneumoniaMNIST and PathMNIST its compact encoder surpasses a from-scratch ViT-B/16 (<InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(\sim \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>∼</mo> </math></EquationSource> </InlineEquation>86&#xa0;M parameters) in macro-F1 while using under 1,000 trainable parameters. The architecture operates entirely on grayscale inputs, confirming its reliance on morphological semantics rather than color bias.</p> Conclusions <p>H-QDCT demonstrates that quantum spectral feature extraction achieves clinically meaningful classification with orders-of-magnitude parameter reduction. By mitigating the convergence instability inherent in small-scale classical transformers, H-QDCT establishes a robust design principle for compact, frequency-aware medical diagnostics on near-term quantum hardware.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

H-QDCT: hierarchical quantum DCT for structural–textural feature fusion in medical imaging

  • Xavier Font Aragones,
  • Miguel Ángel González Ballester

摘要

Purpose

Deep learning models for medical image classification typically rely on millions of parameters or extensive pretraining, limiting deployment in resource-constrained settings. We investigate whether a hybrid quantum-classical architecture can achieve competitive diagnostic performance with extreme parameter efficiency and superior stability compared to lightweight classical counterparts.

Methods

We propose the hierarchical quantum discrete cosine transform (H-QDCT), a lightweight pipeline combining classical patch-based preprocessing with quantum frequency-domain feature extraction. Images are partitioned into patches, amplitude-embedded into quantum states, and transformed via a quantum DCT that reorganizes information by spatial frequency. A hierarchical variational ansatz processes structural (low-frequency) and textural (high-frequency) components in disjoint qubits subspaces before global fusion. We evaluate H-QDCT on six clinically binarized MedMNIST datasets under grayscale constraints against both heavy and lightweight classical baselines.

Results

H-QDCT achieves competitive performance with only 1726 parameters, exceeding a 99.98% reduction compared to ResNet-18 (11.2M parameters). On PneumoniaMNIST, the model attains an AUC of 0.90, approaching the 0.93 of pretrained ResNet-18. Crucially, H-QDCT demonstrates superior robustness in data-scarce regimes: While the lightweight vision transformer (Tiny-ViT) collapsed to the majority class on DermaMNIST (F1 \(=0.00\) = 0.00 ) and BloodMNIST (F1 \(=0.06\) = 0.06 ), H-QDCT retained learnability on BloodMNIST (F1 up to 0.22) with stable convergence. Beyond binary screening, the architecture extends natively to multi-class differential diagnosis (from \(K=2\) K = 2 to \(K=9\) K = 9 ) without mode collapse, and on PneumoniaMNIST and PathMNIST its compact encoder surpasses a from-scratch ViT-B/16 ( \(\sim \) 86 M parameters) in macro-F1 while using under 1,000 trainable parameters. The architecture operates entirely on grayscale inputs, confirming its reliance on morphological semantics rather than color bias.

Conclusions

H-QDCT demonstrates that quantum spectral feature extraction achieves clinically meaningful classification with orders-of-magnitude parameter reduction. By mitigating the convergence instability inherent in small-scale classical transformers, H-QDCT establishes a robust design principle for compact, frequency-aware medical diagnostics on near-term quantum hardware.