<p>The integration of machine learning (ML) models enhances the efficiency, affordability, and reliability of feature detection in microscopy, yet their development and applicability are hindered by the dependency on scarce and often flawed manually labeled datasets with a lack of domain awareness. We addressed these challenges by creating a physics-based synthetic image and data generator, resulting in an ML model that achieves comparable precision (0.86), recall (0.63), F1 scores (0.71), and engineering property predictions (<i>R</i><sup>2</sup> = 0.82) to a model trained on human-labeled data. We enhanced both models by using feature prediction confidence scores to derive an image-wide confidence metric, enabling simple thresholding to eliminate ambiguous and out-of-domain images, resulting in performance boosts of 5–30% with a filtering-out rate of 25%. Our study demonstrates that synthetic data can eliminate human reliance in ML and provides a means for domain awareness in cases where many feature detections per image are needed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accelerating domain-aware electron microscopy analysis using deep learning models with synthetic data and image-wide confidence scoring

  • M. J. Lynch,
  • R. Jacobs,
  • G. A. Bruno,
  • P. Patki,
  • D. Morgan,
  • K. G. Field

摘要

The integration of machine learning (ML) models enhances the efficiency, affordability, and reliability of feature detection in microscopy, yet their development and applicability are hindered by the dependency on scarce and often flawed manually labeled datasets with a lack of domain awareness. We addressed these challenges by creating a physics-based synthetic image and data generator, resulting in an ML model that achieves comparable precision (0.86), recall (0.63), F1 scores (0.71), and engineering property predictions (R2 = 0.82) to a model trained on human-labeled data. We enhanced both models by using feature prediction confidence scores to derive an image-wide confidence metric, enabling simple thresholding to eliminate ambiguous and out-of-domain images, resulting in performance boosts of 5–30% with a filtering-out rate of 25%. Our study demonstrates that synthetic data can eliminate human reliance in ML and provides a means for domain awareness in cases where many feature detections per image are needed.