QdGF: A Framework Which Leverages Training-Time Information for Enhanced Multilingual-Multitask Classification
摘要
We introduce QdGF, a question-driven generative framework that leverages supplementary training data as instructional prompts to enhance multilingual, multitask text classification across multiple generative models. By reframing additional tasks as questions derived from fine-tuning data, QdGF improves cross-lingual generalization while reducing reliance on in-domain pretraining. Our method closes the performance gap between model sizes and tokenization strategies, achieving absolute improvements of +10.5%, +8.37%, and +8.23% in weighted multilingual accuracy across three classification problems using different models. Additionally, QdGF proves particularly effective for underrepresented languages, yielding +19% in Swedish, +37.65% and +28.36% in Hungarian for each of the three classification problems across different models, and for imbalanced datasets, where it increases F1 score by 90% for the most underrepresented class in Dutch. Therefore, QdGF is able to improve classification in data-scarce settings effectively replacing the need for complex pretraining strategies or high-resource models, and regardless of the architecture, tokenization method, or model size.