Abstract <p>Machine learning (ML) methods are increasingly applied to particle identification (PID). Since experimental data lack ground-truth labels, classifiers are trained on Monte Carlo (MC) simulations, which leads to the problem of data shift–distributional differences between simulated and real data. This study analyzes the impact of data shift by comparing particle classification across several MC datasets with different configurations, highlighting the importance of validating and adapting ML models for robust performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Shift Problem in Particle Identification

  • V. V. Papoyan

摘要

Abstract

Machine learning (ML) methods are increasingly applied to particle identification (PID). Since experimental data lack ground-truth labels, classifiers are trained on Monte Carlo (MC) simulations, which leads to the problem of data shift–distributional differences between simulated and real data. This study analyzes the impact of data shift by comparing particle classification across several MC datasets with different configurations, highlighting the importance of validating and adapting ML models for robust performance.