Fine-Tuning Small Language Models for Domain-Specific AI: An Edge AI Perspective
摘要
Deploying large-scale language models on edge devices faces inherent challenges such as high computational demands, energy consumption, and potential data privacy risks. This study introduces the Shakti Small Language Models (SLMs)—Shakti-100M, Shakti-250M, and Shakti-500M—which target these constraints head-on. By combining efficient architectures, quantization techniques, and responsible AI principles, the Shakti series enables on-device intelligence for smartphones, smart appliances, IoT systems, and beyond. We provide comprehensive insights into their design philosophy, training pipelines, and benchmark performance on both general tasks (e.g. MMLU, HellaSwag, where Shakti-500M achieves 65.3% and 48.7% accuracy, respectively) and do-main-specific tasks such as healthcare, finance, and legal applications, where Shakti-250M demonstrates competitive performance against larger models. Our findings illustrate that compact models, when carefully engineered and fine-tuned, can meet and often exceed expectations in real-world domain-specific edge-AI scenarios.