Tree-Based Classification of the Technical Ukrainian Texts
摘要
Our research leverages anonymous data from an online garage management system. This vast dataset, enriched by transcribed customer calls, forms the basis for our tree-based classification model development, aimed at categorizing automotive repairs and parts. Addressing the challenges of analyzing technical texts, particularly in Ukrainian, we adapted classical NLP methods for our dataset. We explored the most popular data classification algorithms. Despite their potential, none provided sufficient accuracy for classifying parts, leading us to develop and train a custom algorithm tailored to our hierarchical data structure. Our Python library facilitates real-time classification, demonstrating significant speed and accuracy improvements across various datasets. This work not only advances the application of tree-based classification in the automotive sector but also sets the stage for future research on more complex tasks, including extracting and classifying automotive information from large text corpora. In conclusion, our study underscores the utility of tree-based classification for technical text analysis within the automotive industry, promising enhanced operational efficiencies and valuable insights for various stakeholders.