Weighted Logistic Regression Model
摘要
Binary responses, which are common in surveys, are modeled through binary models, a relationship between the probability of the response and a set of covariates. However, as explained in Chap. 4 , when the observations are not obtained by simple random sampling, the standard logistic regression is not valid. When the data come from a complex survey designed with stratification, clustering, and/or unequal weighting, the usual estimates are not appropriate (Rao & Scott, 1984). In these cases, specialized techniques are applied in order to produce the appropriate estimates and their standard errors. Clustered data are frequently encountered in fields such as health services, public health, epidemiology, and education research. Data may consist of patients clustered within primary care practices or hospitals, or households clustered within neighborhoods, or of students clustered within schools. Subjects nested within the same cluster often exhibit a greater degree of similarity, or homogeneity of outcomes, compared to randomly selected subjects from different clusters (Multilevel analysis: an introduction to basic and advanced multilevel modeling, Thousand Oaks, CA; Hierarchical linear models: applications and data analysis methods, Thousand Oaks, CA; Introduction to multilevel modeling, Thousand Oaks, CA; Multilevel statistical models, London; Canadian Journal of Public Health 92:150–154, 2001). Due to the possible lack of independence of subjects within the same cluster, traditional statistical methods are not appropriate for the analysis of clustered data. While Chap. 4 , uses the overdispersed logistic regression and the exchangeability logistic regression model to fit correlated data, this chapter incorporates a series of weights or design effects to account for the correlation. The logistic regression on the analysis of survey data takes into account the properties of the survey sample design, including stratification, clustering, and unequal weighting. The chapter fits this model in SAS, SPSS, and R, using methods based on the following: