<p>This study investigates the consistency of statistical tests and the effectiveness of effect size measures in evaluating items with Differential Item Functioning (DIF) in a large-scale assessment. We applied a stepwise binary Logistic Regression (LR) analysis to a large-scale exam sat by 999 PhD applicants in Iran. First, we employed Likelihood Ratio Test (LRT) and Wald separately on the same dichotomous data to detect their equivalence in identifying DIF. The LRT compares two models to see whether including group membership (e.g., gender) significantly improves the model's fit to the data, while the Wald test checks whether the difference in performance between groups is statistically significant based on model coefficients. Then to evaluate the effectiveness of effect size measures, we combined Nagelkerke Δ<sub>R</sub><sup>2</sup> with statistical tests and, in a separate analysis, blended Δ log of odds ratio <b>(</b>Δ <sub>LR</sub><b>)</b> effect size criterion with confidence intervals. Δ<sub>R</sub><sup>2</sup> estimates how much of the variation in responses is explained by group membership (similar to a percentage of explained variance), and D<sub>LR</sub> quantifies how much more or less likely one group is to answer an item correctly compared to another group. The results indicated that blending the effect size measures with statistical tests or confidence intervals was more effective in identifying practically significant DIF-flagged items than using the statistical tests alone. Furthermore, it was indicated that Δ <sub>LR</sub> was less conservative than Δ <sub>R</sub><sup>2</sup> in identifying practically significant DIF-flagged items. We suggest that assessment developers apply effect size measures combined with statistical tests to identify practically significant DIF-flagged items, hence ensuring assessment fairness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating differential item functioning through effect size measures in logistic regression: implications for language assessment

  • Ali Darabi Bazvand,
  • Balal Izanloo,
  • Seyed Mohammad Reza Amirian,
  • Kaveh Jalilzadeh

摘要

This study investigates the consistency of statistical tests and the effectiveness of effect size measures in evaluating items with Differential Item Functioning (DIF) in a large-scale assessment. We applied a stepwise binary Logistic Regression (LR) analysis to a large-scale exam sat by 999 PhD applicants in Iran. First, we employed Likelihood Ratio Test (LRT) and Wald separately on the same dichotomous data to detect their equivalence in identifying DIF. The LRT compares two models to see whether including group membership (e.g., gender) significantly improves the model's fit to the data, while the Wald test checks whether the difference in performance between groups is statistically significant based on model coefficients. Then to evaluate the effectiveness of effect size measures, we combined Nagelkerke ΔR2 with statistical tests and, in a separate analysis, blended Δ log of odds ratio (Δ LR) effect size criterion with confidence intervals. ΔR2 estimates how much of the variation in responses is explained by group membership (similar to a percentage of explained variance), and DLR quantifies how much more or less likely one group is to answer an item correctly compared to another group. The results indicated that blending the effect size measures with statistical tests or confidence intervals was more effective in identifying practically significant DIF-flagged items than using the statistical tests alone. Furthermore, it was indicated that Δ LR was less conservative than Δ R2 in identifying practically significant DIF-flagged items. We suggest that assessment developers apply effect size measures combined with statistical tests to identify practically significant DIF-flagged items, hence ensuring assessment fairness.