Background <p>Asthma is a growing public health concern in India, but research has largely focused on children or general adult populations, often overlooking women of reproductive age. Prior studies typically use linear models that fail to capture the complex interactions among environmental, socio-demographic, and behavioural risk factors. This study addresses this gap by estimating asthma prevalence in Indian women (15–49 years) and applying machine learning techniques to identify non-linear, high-dimensional predictors using NFHS-5 data.</p> Methods <p>This study analysed NFHS-5 data (2019–2021) using a nationally representative stratified two-stage sampling design. A total of 550,746 women aged 15–49 was included after excluding non-responses to asthma-related questions. Asthma status was self-reported. Bivariate Chi-square tests examined associations with environmental, socio-economic, behavioral, nutritional, and geographic variables. A one-sample t-test assessed dietary score differences. Three machine learning models like Logistic Regression, Random Forest, and XGBoost were developed on a balanced dataset using up-sampling in R (caret package). Model performance was evaluated using AUC and accuracy; key predictors were identified via feature importance and predicted probabilities.</p> Results <p>Asthma prevalence was 15.4 per 1,000 women (95% CI: 14.9–16.0). Significant associations were observed with environmental (housing, fuel type, sanitation), socio-economic (age, education, caste, religion), behavioral (tobacco, alcohol), and nutritional factors (BMI, dietary score). Random Forest outperformed other models (AUC: 0.912; accuracy: 84.3%; <i>p</i> &lt; 0.001), with dietary score, age, and wealth index as top predictors. Predicted risk was notably higher among older, overweight, less-educated, and urban women (<i>p</i> &lt; 0.05 for all comparisons).</p> Conclusions <p>Asthma among Indian women is shaped by diverse social, environmental, and behavioral factors. Machine learning, particularly Random Forest, offers valuable predictive insights, highlighting high-risk groups and supporting targeted public health interventions.</p> Trial registration <p>Not applicable.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prevalence and predictors of asthma among Indian women: a machine learning-based analysis of NFHS-5 data

  • Vini Mehta,
  • Anil Pardeshi,
  • Rayhan Rahman,
  • Illias Sheikh,
  • Ankita Mathur

摘要

Background

Asthma is a growing public health concern in India, but research has largely focused on children or general adult populations, often overlooking women of reproductive age. Prior studies typically use linear models that fail to capture the complex interactions among environmental, socio-demographic, and behavioural risk factors. This study addresses this gap by estimating asthma prevalence in Indian women (15–49 years) and applying machine learning techniques to identify non-linear, high-dimensional predictors using NFHS-5 data.

Methods

This study analysed NFHS-5 data (2019–2021) using a nationally representative stratified two-stage sampling design. A total of 550,746 women aged 15–49 was included after excluding non-responses to asthma-related questions. Asthma status was self-reported. Bivariate Chi-square tests examined associations with environmental, socio-economic, behavioral, nutritional, and geographic variables. A one-sample t-test assessed dietary score differences. Three machine learning models like Logistic Regression, Random Forest, and XGBoost were developed on a balanced dataset using up-sampling in R (caret package). Model performance was evaluated using AUC and accuracy; key predictors were identified via feature importance and predicted probabilities.

Results

Asthma prevalence was 15.4 per 1,000 women (95% CI: 14.9–16.0). Significant associations were observed with environmental (housing, fuel type, sanitation), socio-economic (age, education, caste, religion), behavioral (tobacco, alcohol), and nutritional factors (BMI, dietary score). Random Forest outperformed other models (AUC: 0.912; accuracy: 84.3%; p < 0.001), with dietary score, age, and wealth index as top predictors. Predicted risk was notably higher among older, overweight, less-educated, and urban women (p < 0.05 for all comparisons).

Conclusions

Asthma among Indian women is shaped by diverse social, environmental, and behavioral factors. Machine learning, particularly Random Forest, offers valuable predictive insights, highlighting high-risk groups and supporting targeted public health interventions.

Trial registration

Not applicable.