<p>Accurate loss cause coding is foundational to insurance pricing, reserving, exposure management, and actuarial data integrity. This study evaluates whether claim narrative text can identify narrative–code inconsistencies as proxy signals for potential loss cause misclassification requiring expert review. We propose a production-oriented hierarchical classification framework in which a first-stage router predicts a broad peril category and routes each claim to a peril-specific submodel for detailed cause classification. Rule-based baselines, TF-IDF classical machine learning models, and text-embedding classifiers were evaluated across top-level and within-peril tasks. On a pass-filtered top-level dataset comprising 1,080 claims across nine peril classes, embedding-based logistic regression achieved 0.9704 accuracy and 0.9696 macro F1. Within-peril embedding models achieved macro F1 scores of 0.9117 for winter weather, 0.9063 for internal water, and 0.8757 for flood/hurricane. Full hierarchical pipeline evaluation over 511 held-out claims reveals an 8.6%-point oracle-to-pipeline accuracy gap (0.9119 vs. 0.8258), demonstrating that hierarchical ML systems must be evaluated as integrated pipelines rather than isolated components; because the top-level router and within-peril submodels were trained on differently filtered corpora (§&#xa0;5.2, §&#xa0;12), this gap should be interpreted as an upper-bound estimate of routing error propagation rather than a fully isolated, confound-free measurement. A confidence threshold analysis identifies operating points translating model confidence into operational review volume and error capture rate. A preliminary adjudication pilot of all 159 low-confidence flagged claims (confidence below 0.80) by a domain expert yielded an adjudicated flag precision of 31.7% (Wilson 95% CI: 19.6%–47.0%) and a useful-flag rate of 30.1% (Wilson 95% CI: 23.5%–37.7%), representing an approximate 3.2× enrichment of correctable claims relative to random review; 34 of 156 adjudicated records (21.8%) had both the system-assigned code and model prediction independently found to be incorrect, demonstrating that confidence-threshold flagging surfaces claims requiring expert judgment beyond what either automated source captures. The framework produces a prioritized human review queue rather than automatic corrections, preserving auditability and regulatory compliance in regulated insurance environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A hierarchical machine learning framework for detecting loss cause misclassification in insurance claims using narrative text

  • Naveen Karakavalasa

摘要

Accurate loss cause coding is foundational to insurance pricing, reserving, exposure management, and actuarial data integrity. This study evaluates whether claim narrative text can identify narrative–code inconsistencies as proxy signals for potential loss cause misclassification requiring expert review. We propose a production-oriented hierarchical classification framework in which a first-stage router predicts a broad peril category and routes each claim to a peril-specific submodel for detailed cause classification. Rule-based baselines, TF-IDF classical machine learning models, and text-embedding classifiers were evaluated across top-level and within-peril tasks. On a pass-filtered top-level dataset comprising 1,080 claims across nine peril classes, embedding-based logistic regression achieved 0.9704 accuracy and 0.9696 macro F1. Within-peril embedding models achieved macro F1 scores of 0.9117 for winter weather, 0.9063 for internal water, and 0.8757 for flood/hurricane. Full hierarchical pipeline evaluation over 511 held-out claims reveals an 8.6%-point oracle-to-pipeline accuracy gap (0.9119 vs. 0.8258), demonstrating that hierarchical ML systems must be evaluated as integrated pipelines rather than isolated components; because the top-level router and within-peril submodels were trained on differently filtered corpora (§ 5.2, § 12), this gap should be interpreted as an upper-bound estimate of routing error propagation rather than a fully isolated, confound-free measurement. A confidence threshold analysis identifies operating points translating model confidence into operational review volume and error capture rate. A preliminary adjudication pilot of all 159 low-confidence flagged claims (confidence below 0.80) by a domain expert yielded an adjudicated flag precision of 31.7% (Wilson 95% CI: 19.6%–47.0%) and a useful-flag rate of 30.1% (Wilson 95% CI: 23.5%–37.7%), representing an approximate 3.2× enrichment of correctable claims relative to random review; 34 of 156 adjudicated records (21.8%) had both the system-assigned code and model prediction independently found to be incorrect, demonstrating that confidence-threshold flagging surfaces claims requiring expert judgment beyond what either automated source captures. The framework produces a prioritized human review queue rather than automatic corrections, preserving auditability and regulatory compliance in regulated insurance environments.