<p>Cluster analysis is a useful technique for discovery of natural groups in multivariate data sets, and Ward’s minimum variance clustering method has been one of the most successful. Advances in the use of semi-metric, non-Euclidean distance measures for non-parametric ANOVA (PERMANOVA; Anderson, Wiley statsref: statistics reference online, Wiley, United States, 2014) have justified the use of distance measures such as the Bray–Curtis index with Ward’s clustering algorithm. Experimental use of Ward’s method with Bray–Curtis distance has shown it to perform well relative to other methods. Ward’s method can produce classifications with extremely high and rare PERMANOVA pseudo <i>f</i>-ratios, but they can be improved upon by trial-and-error optimization using the f-ratio as the criterion for re-classifying samples. Thus Ward’s method with Bray–Curtis or other non-Euclidean distance measures is promising for data exploration and as input for methods that optimize an input classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Minimum variance clustering of plant community data with non-Euclidean distance measures: Ward’s method and PERMANOVA

  • David K. Swanson

摘要

Cluster analysis is a useful technique for discovery of natural groups in multivariate data sets, and Ward’s minimum variance clustering method has been one of the most successful. Advances in the use of semi-metric, non-Euclidean distance measures for non-parametric ANOVA (PERMANOVA; Anderson, Wiley statsref: statistics reference online, Wiley, United States, 2014) have justified the use of distance measures such as the Bray–Curtis index with Ward’s clustering algorithm. Experimental use of Ward’s method with Bray–Curtis distance has shown it to perform well relative to other methods. Ward’s method can produce classifications with extremely high and rare PERMANOVA pseudo f-ratios, but they can be improved upon by trial-and-error optimization using the f-ratio as the criterion for re-classifying samples. Thus Ward’s method with Bray–Curtis or other non-Euclidean distance measures is promising for data exploration and as input for methods that optimize an input classification.