Cross-Domain Aspect Sentiment Triplet Extraction Based on Generative Data Augmentation and Pseudo-Label Optimization
摘要
Cross-domain aspect sentiment triplet extraction (CD-ASTE) aims to extract sentiment triplets from unlabeled target domains by transferring knowledge from labeled source domains. However, existing CD-ASTE methods still suffer from insufficient augmented samples and semantic misalignment in pseudo-labels. To address these issues, we propose a multi-stage CD-ASTE framework that integrates extraction and generation tasks. Specifically, we employ paraphrasing-based augmentation to generate diverse target-domain texts to alleviate data sparsity. In addition, we introduce a pseudo-label optimization mechanism to filter noisy pseudo-labels and reduce semantic errors, effectively mitigating error propagation. By combining multi-stage augmentation and pseudo-label filtering, our approach fully exploits unlabeled target-domain data and enhances cross-domain transferability. Experiments on six cross-domain settings show that our model achieves new state-of-the-art (SOTA) performance compared to strong baselines.