What matters in patent claims: structural and semantic-based feature extraction and embedding optimization
摘要
This study introduces HI-ICD-AugCCE, a hybrid framework that integrates structural and semantic analysis of patent claims using deep learning techniques. By assigning importance weights to individual claims based on hierarchical position and information density, the framework effectively captures the core technical content within patent texts. It further employs an unsupervised contrastive learning model enhanced with data augmentation, enabling accurate similarity evaluation even in the absence of labeled data. Experiments demonstrate that HI-ICD-AugCCE achieves a Spearman correlation of 0.812, representing a 6.98% improvement over baselines. The results highlight the framework’s effectiveness in enhancing the semantic precision and scalability of patent similarity assessment.