A Framework for Cross-dataset Semi-supervised Segmentation of Liver Hepatocellular Carcinoma in Whole Slide Images
摘要
Hepatocellular Carcinoma (HCC) is a type of liver cancer that contributes for about 75% of all liver cancers. It is one of the most prevalent primary liver cancers. Nowadays, histopathology tissue analysis has become a leading standard for the accurate diagnosis and treatment of cancer. With the advent of artificial intelligence, computer-assisted diagnosis (CAD) has come to aid histopathologists in whole slide image (WSI) analysis. However, the size of WSIs, variability in the staining protocol and lack of sufficient labelled dataset are some of the most important obstacles encountered by the researchers working in this area of WSI analysis. The paper proposes a framework for cross-dataset learning of segmentation annotations for HCC using a contrastive semi-supervised learning algorithm to address these challenges. This framework uses a pre-labelled dataset, PAIP 2019, to annotate diagnostic slide images from the TCGA-LIHC dataset. Here, cellular patches of TCGA-LIHC are restained using an unpaired image-to-image translation model, CycleGAN, to minimise the superficial differences between the datasets. Pseudo-labelling semi-supervised algorithm is used to train a modified UNet model on the labelled and restained unlabelled images to build a more generalised segmentation framework. The segmentation model trained using the proposed framework performs better than the model trained using the labelled dataset only.