Heterogeneity-aware orthogonal covariance attention for MRI-based multi-class brain tumor differentiation
摘要
Non-invasive, accurate classification of brain tumours via MRI is a crucial foundation for clinical diagnosis and personalised treatment decisions. However, existing deep learning models generally rely on global average pooling for feature aggregation, this mechanism distorts spatial prior information about the lesion and masks the high internal heterogeneity of tumours, thereby limiting the models’ discriminative power and clinical reliability. This paper proposed the Orthogonal Covariance Attention Module (OCAM), which reconstructs channel-spatial covariance through low-rank second-order statistical modelling and achieves adaptive feature compression by combining a mask-guided weighted aggregation mechanism. Building upon this, an orthogonal regularisation constraint is introduced to decouple the feature basis, enabling the model to perform weakly supervised lesion localisation relying solely on image-level labels. This module unifies the modelling of spatial heterogeneity and higher-order feature relationships with linear computational complexity and can be seamlessly integrated into both CNN and Transformer architectures. On multi-class brain tumour MRI datasets, OCAM achieves consistent performance improvements across different backbone networks and training strategies. When applied to a pre-trained ResNet18, the overall accuracy is improved to 96.06% and specificity reaches 98.69%, outperforming several mainstream feature aggregation methods. Experiments across ResNet50 and Swin-Transformer further validate its strong generalisation capability. Visualisation results demonstrate that OCAM can significantly suppress background noise and precisely focus on lesion regions, whilst achieving an effective balance between accuracy and inference efficiency with virtually no increase in parameter count or computational complexity. OCAM overcomes the expressive limitations of traditional discriminative feature aggregation at the level of statistical feature modelling, achieving structured modelling of the spatial heterogeneity of brain tumours and the decoupling of discriminative features. This not only enhances classification performance but also significantly improves the model’s interpretability and clinical credibility. Whilst maintaining a lightweight computational overhead, this method possesses cross-architecture adaptability and practical deployment potential, providing a universal and implementable solution for highly reliable medical image-assisted diagnosis systems.