Data Shift Problem in Particle Identification
摘要
Abstract
Machine learning (ML) methods are increasingly applied to particle identification (PID). Since experimental data lack ground-truth labels, classifiers are trained on Monte Carlo (MC) simulations, which leads to the problem of data shift–distributional differences between simulated and real data. This study analyzes the impact of data shift by comparing particle classification across several MC datasets with different configurations, highlighting the importance of validating and adapting ML models for robust performance.