Misclassification
摘要
Epidemiological analyses are often based on categorical variables and often such variables are prone to misclassification, i.e., for some study subjects, the recorded and actual values of the variable differ. For instance, some study subjects may respond erroneously when asked to self-report a categorial variable on a questionnaire. This chapter first outlines the deleterious impacts of ignoring misclassification, i.e., treating the recorded data as if it were fully correct. Then various strategies for acknowledging misclassification in the analysis are reviewed. Any such strategy requires information about the magnitude of the misclassification. In the case of a binary variable, for instance, that could be information about the sensitivity and specificity of the observable variable, as a surrogate for the actual, but latent, variable of interest. In keeping with previous literature, most attention is paid to the situation that a categorical exposure variable is misclassified. However, misclassification in a confounding variable, and in an outcome variable, is also discussed.