<p>Untargeted metabolomics is particularly appropriate for the identification of new biomarkers of different diseases. The approach requires analysis of high-dimensional datasets, which is impossible without automated data preprocessing. The study shows possibilities of various software for data preprocessing, i.e., MZmine, XCMS, MS-DIAL, iMet-Q, and Peakonly, to analyze the results of untargeted urine analysis of 40 cancer patients and 40 healthy volunteers using ultra high-performance liquid chromatography coupled with high resolution mass spectrometry (UHPLC-HRMS). The programs were compared based on peak integration quality, feature selection in untargeted analysis, accuracy of models created using those features, and statistical analysis of peaks of interest. Mann–Whitney <i>U</i> test and random forest were used to identify the list of the most important features, which were then used as input data for creating diagnostic models using logistic regression. All software tools show different features as most important. Suberoyl-<span>l</span>-carnitine and Tetranor-PGJM were selected as most important in MZmine and XCMS. Data preprocessing tools significantly affect the results of statistical analysis. Peak integration was comparable to manual in case of MS-DIAL, MZmine, and Peakonly. MS-DIAL and XCMS are preferable for identification. iMet-Q gives good results for statistical analysis, performance of diagnostic models, and finding compounds of interest. MS-DIAL and iMet-Q were preferable for both identification and statistical analysis making them desirable tools for untargeted metabolomics in clinical studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparison of various data preprocessing software for untargeted urinary metabolomics using UHPLC-HRMS to diagnose cancer

  • Elina Gashimova,
  • Tatyana Malitskaya,
  • Azamat Temerdashev,
  • Dmitry Perunov,
  • Vladimir Porkhanov,
  • Igor Polyakov

摘要

Untargeted metabolomics is particularly appropriate for the identification of new biomarkers of different diseases. The approach requires analysis of high-dimensional datasets, which is impossible without automated data preprocessing. The study shows possibilities of various software for data preprocessing, i.e., MZmine, XCMS, MS-DIAL, iMet-Q, and Peakonly, to analyze the results of untargeted urine analysis of 40 cancer patients and 40 healthy volunteers using ultra high-performance liquid chromatography coupled with high resolution mass spectrometry (UHPLC-HRMS). The programs were compared based on peak integration quality, feature selection in untargeted analysis, accuracy of models created using those features, and statistical analysis of peaks of interest. Mann–Whitney U test and random forest were used to identify the list of the most important features, which were then used as input data for creating diagnostic models using logistic regression. All software tools show different features as most important. Suberoyl-l-carnitine and Tetranor-PGJM were selected as most important in MZmine and XCMS. Data preprocessing tools significantly affect the results of statistical analysis. Peak integration was comparable to manual in case of MS-DIAL, MZmine, and Peakonly. MS-DIAL and XCMS are preferable for identification. iMet-Q gives good results for statistical analysis, performance of diagnostic models, and finding compounds of interest. MS-DIAL and iMet-Q were preferable for both identification and statistical analysis making them desirable tools for untargeted metabolomics in clinical studies.