Compact Quantum Modules for Scalable Fine-Tuning in Large Language Models
摘要
The rapid growth of large language models (LLMs) has intensified the demand for efficient fine-tuning techniques under constrained computational and memory budgets. Parameter-efficient fine-tuning (PEFT) addresses this challenge, yet existing methods often sacrifice expressive capacity for reduced complexity. This paper introduces Quantum-amplitude embedded (QAE) adaptation, a novel PEFT framework that leverages quantum-amplitude embedding to achieve logarithmic compression of activation vectors while preserving task-relevant information. By integrating parameterized quantum circuits (PQCs) as nonlinear transformation layers, QAE adaptation replaces conventional linear adapters in attention modules with compact quantum-inspired operators. This design provides high expressivity with significantly fewer trainable parameters. Experimental evaluations across benchmark tasks demonstrate that QAE achieves competitive or superior performance compared with state-of-the-art PEFT methods, which highlights its effectiveness for resource-constrained LLM adaptation.