Some Robust Data Analyses
摘要
In this last chapter, we exhibit the power of the methods described earlier, by analysing five datasets. We start in Sect. 10.2 with the two sets of income data from Sect. 1.4 . Without explanatory variables, we found the log transformationTransformation for the former, while the analysis of the latter remained inconclusive. When explanatory variables are included, and outliersOutlier deleted, the square-root transformationTransformation is indicated for both. In Sect. 10.4 we analyse 1711 responses to a survey on customer loyalty, in which there are six explanatory variables. Parametric methods lead to \(\surd {y}\) as the response, the identification of 41 outliersOutlier, and a skewed distribution of residualsResiduals. RAVAS followed by the FSFS provides a good approximation to normally distributed errors, when only nine observations are deleted. This analysis is summarized in tabular form in Sect. 10.4.6 to provide a template for the modern robust analysis of regression data. Despite transformationTransformation and outlierOutlier detection, the t-statistics statistics for the significance of the variables in the customer loyalty dataLoyalty data hardly change. Accordingly, in Sect. 10.5 we modify 25 observations: monitoring plots reveal the outliersOutlier and the results of the RAVAS analysis are close to those for the uncontaminated data. Finally, we analyse the NCI-60 cancer cell data (Chap. 9 ). With only seven explanatory variables, we monitor LS diagnostics, detect outliersOutlier, and apply RAVAS, which gives the best-fitting model. The generalized candlestick plot provides further model selection.