Background <p>Radiotracer labels in positron emission tomography (PET) DICOM metadata are manually entered and often unreliable or missing, requiring expert visual verification for downstream analyses. We developed and validated deep learning models to classify PET tracers directly from image data across multiple sites and cohorts.</p> Results <p>We used PET scans obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI, <i>n</i> = 7429) and the Standardized Centralized Alzheimer’s &amp; Related Dementias Neuroimaging study (SCAN, <i>n</i> = 4869) to classify 10 tracers. We evaluated generalizability on an external test set of Mayo Clinic Study of Aging and ADRC scans containing four tracers (<i>n</i> = 10558). ADNI and SCAN data were combined and split into training (60%), validation (20%), and internal test (20%) sets. We trained two ConvNeXt-Large models initialized with ImageNet weights: a 10-class model (AV1451, FBP, FBP-Early, FBB, FBB-Early, FDG, MK6240, NAV4694, PI2620, PIB) and a 3-class model grouping tracers by biological target (Amyloid, FDG, Tau). Performance was evaluated using macro-F1, average precision (AP), and confusion matrices, Integrated Gradients for attribution, with Uniform Manifold Approximation and Projection (UMAP) and cosine similarity to assess model interpretability and multi-cohort effects on the embedding structure. The 3-class model achieved macro-F1 ≥ 0.99 across all test sets. The 10-class model achieved macro-F1 of 0.91 on the internal test and 0.93 on the external test set where only four tracers were available. Lower performance was observed for NAV4694 (AP 0.70; misclassified as FBB, 17.2%), PI2620 (0.79; AV1451, 20.0%), and FBB (0.80; FBP, 23.3%). The attribution maps indicated the models learned some tracer specific uptake patterns as well as general imaging features. Class embeddings in the UMAP projection showed clear separation, with limited overlap among underrepresented tracers with similar binding characteristics. Cohort-level differences in learned representations were most pronounced among tau tracers, with higher within-class cosine similarity in ADNI than in SCAN (mean [SD], 0.26 [0.10] vs. 0.01 [0.06]), while amyloid tracers showed minimal shift across cohorts.</p> Conclusions <p>Classification was near-perfect when grouping tracers into three biological targets, and the 10-class model generalized well to the four tracers available in the independent external cohort. A few tracers with similar binding characteristics showed some overlap in latent space, with minimal impact on overall classification performance. Our results demonstrate that deep learning models can identify PET tracers from image appearance alone and have the potential to support semi-automated tracer identification and data curation in large multi-site PET studies. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated PET tracer classification in multi-site, multi-cohort studies using deep learning

  • Robel K. Gebre,
  • Robert I. Reid,
  • Matthew L. Senjem,
  • Jeffrey L. Gunter,
  • Val J. Lowe,
  • Kejal Kantarci,
  • Ronald C. Petersen,
  • Jonathan Graff-Radford,
  • Clifford R. Jack Jr.,
  • Prashanthi Vemuri,
  • Christopher G. Schwarz

摘要

Background

Radiotracer labels in positron emission tomography (PET) DICOM metadata are manually entered and often unreliable or missing, requiring expert visual verification for downstream analyses. We developed and validated deep learning models to classify PET tracers directly from image data across multiple sites and cohorts.

Results

We used PET scans obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI, n = 7429) and the Standardized Centralized Alzheimer’s & Related Dementias Neuroimaging study (SCAN, n = 4869) to classify 10 tracers. We evaluated generalizability on an external test set of Mayo Clinic Study of Aging and ADRC scans containing four tracers (n = 10558). ADNI and SCAN data were combined and split into training (60%), validation (20%), and internal test (20%) sets. We trained two ConvNeXt-Large models initialized with ImageNet weights: a 10-class model (AV1451, FBP, FBP-Early, FBB, FBB-Early, FDG, MK6240, NAV4694, PI2620, PIB) and a 3-class model grouping tracers by biological target (Amyloid, FDG, Tau). Performance was evaluated using macro-F1, average precision (AP), and confusion matrices, Integrated Gradients for attribution, with Uniform Manifold Approximation and Projection (UMAP) and cosine similarity to assess model interpretability and multi-cohort effects on the embedding structure. The 3-class model achieved macro-F1 ≥ 0.99 across all test sets. The 10-class model achieved macro-F1 of 0.91 on the internal test and 0.93 on the external test set where only four tracers were available. Lower performance was observed for NAV4694 (AP 0.70; misclassified as FBB, 17.2%), PI2620 (0.79; AV1451, 20.0%), and FBB (0.80; FBP, 23.3%). The attribution maps indicated the models learned some tracer specific uptake patterns as well as general imaging features. Class embeddings in the UMAP projection showed clear separation, with limited overlap among underrepresented tracers with similar binding characteristics. Cohort-level differences in learned representations were most pronounced among tau tracers, with higher within-class cosine similarity in ADNI than in SCAN (mean [SD], 0.26 [0.10] vs. 0.01 [0.06]), while amyloid tracers showed minimal shift across cohorts.

Conclusions

Classification was near-perfect when grouping tracers into three biological targets, and the 10-class model generalized well to the four tracers available in the independent external cohort. A few tracers with similar binding characteristics showed some overlap in latent space, with minimal impact on overall classification performance. Our results demonstrate that deep learning models can identify PET tracers from image appearance alone and have the potential to support semi-automated tracer identification and data curation in large multi-site PET studies.