In the realm of official statistics, the categorization of economic activities plays a crucial role in analyzing and interpreting data. The NACE (Nomenclature générale des activités économiques dans les communautés européennes) coding system is widely utilized for this purpose, but manual coding by domain experts is time-consuming and prone to inconsistencies. Thus, there is a growing need to automate the NACE coding process to address these challenges. In the present study, we delineate the methods devised for the automated classification of German language activity descriptors sourced from two statistical entities: namely, Statistics Austria and the Federal Statistical Office of Germany. Using machine learning paradigms alongside natural language processing (NLP) strategies, this investigation examines the distinctive data origins and computational techniques within the domains of each institution. Despite variations in data provenance and algorithmic techniques, our findings reveal a convergence in the encountered challenges and outcomes. Notable is the discernible impact of hierarchical data structure on improving classification results across both scenarios, albeit with varying degrees of efficacy. Of particular interest is the persistent complexity associated with the accurate prediction of NACE codes in both analytical approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Approaches to Automated NACE Coding of German Business Activity Descriptions

  • Felix Beuter,
  • Johannes Gussenbauer,
  • Elias Minther,
  • Viktoria Szabo,
  • Susanne Wegner

摘要

In the realm of official statistics, the categorization of economic activities plays a crucial role in analyzing and interpreting data. The NACE (Nomenclature générale des activités économiques dans les communautés européennes) coding system is widely utilized for this purpose, but manual coding by domain experts is time-consuming and prone to inconsistencies. Thus, there is a growing need to automate the NACE coding process to address these challenges. In the present study, we delineate the methods devised for the automated classification of German language activity descriptors sourced from two statistical entities: namely, Statistics Austria and the Federal Statistical Office of Germany. Using machine learning paradigms alongside natural language processing (NLP) strategies, this investigation examines the distinctive data origins and computational techniques within the domains of each institution. Despite variations in data provenance and algorithmic techniques, our findings reveal a convergence in the encountered challenges and outcomes. Notable is the discernible impact of hierarchical data structure on improving classification results across both scenarios, albeit with varying degrees of efficacy. Of particular interest is the persistent complexity associated with the accurate prediction of NACE codes in both analytical approaches.