Multiple Comparison and Post Hoc Tests

  • 1. Bonferroni-adjusted multiple t-tests (Dunn)
  • 2. Sidak test
  • 3. Dunnett's test
  • 4. Tukey honestly significant difference (HSD) test
  • 5. Games and Howell's modification of Tukey's HSD
  • 6. Tukey's wholly significant difference (WSD) test
  • 7. Newman-Keuls test (Student-Newman-Keuls)
  • 8. Ryan test (REGWQ)
  • 9. The Shaffer-Ryan test
  • 10. The least significant difference test (LSD) / Fisher's LSD
  • 11. The Fisher-Hayter test
  • 12. Waller-Duncan test
  • 13. Games-Howell GH
  • 14. Dunnett's T3 and Dunnett's C
  • 15. Tamhane's T2
  • 16. The Tukey-Kramer test
  • 17. The Miller-Winer test
  • 18. Multiple range (homogeneous subset) tests
  • 19. The Scheffé test
  • 20. The Duncan's test
  • 21. Hochberg's GT2 test
  • 22. Gabriel test

Post Hoc Matrix Parameters

Test Name Property / Approach When to Use Remarks
Bonferroni-adjusted multiple t-tests (Dunn) It is a simple type of multiple comparison test. 1) Used only when there are few numbers of comparisons.
2) It can be suitable for nonpairwise as well as pairwise comparisons.
Baseline direct control.
Sidak test The alpha significance level for multiple comparisons is closer than the Bonferroni test. Post hoc testing distributions. It controls the family-wise error rate when the comparisons are orthogonal to each other.
Dunnett's test Control reference classification. Used when one wants to compare each treatment group mean with the mean of the control group. Target control variance.
Tukey honestly significant difference (HSD) test It is a very conservative pairwise comparison test as large number of groups would inflate Type I errors. 1) When the number of groups is large.
2) When all pairwise comparisons are tested.
Post Hoc Conservative
Games and Howell's modification of Tukey's HSD Modified HSD test variance tracking. Suitable when the homogeneity of variances assumption is violated. Post Hoc relatively liberal.
Tukey's wholly significant difference (WSD) test Less conservative version of Tukey's HSD. Multiple variance ranges. Post Hoc less conservative.
Newman-Keuls test (Student-Newman-Keuls) Based on the q-statistic, which is used to evaluate partial null hypotheses. 1) Recommended when the researcher wants to compare adjacent means.
2) Only when the number of groups to be compared equals three.
Stepwise bounds.
Ryan test (REGWQ) 1) Modified version of Newman-Keuls test where the alpha level decreases when stretch size decreases.
2) Controls alpha rate safely even when groups exceed three with balanced power.
3) Step-down procedure.
Multi-group classification. Best choice as it maintains good alpha control with 75% of the power.
The Shaffer-Ryan test Modified Ryan test and also step-down test. High power comparison analysis. Best multiple comparison tests in terms of power parameters.
The least significant difference test (LSD) / Fisher's LSD Based on the t-statistic metrics. 1) Suitable for both pairwise and nonpairwise comparisons.
2) It will not require equal sample sizes.
LSD is the most liberal of the post-hoc tests with poor control of alpha.
The Fisher-Hayter test Modified LSD with structured control on alpha parameters. Suitable when all pairwise comparisons are done post-hoc, but power may be low for fewer comparisons. Alpha optimized framework.
The Scheffé test Checks whether overall null hypothesis is rejected. If rejected, F values are computed for all possible comparisons. When the number of comparisons is large. Mostly used method to control alpha at the cost of less power. More conservative than Tukey test.
The Duncan's test Identical to Student Newman-Keuls test. It uses error rate for collection of test data. Homogeneous ranges. Liberal range subsets.
Hochberg's GT2 test Similar to Tukey's Honestly Significant test; uses studentized maximum modulus. Unbalanced sample sets. Modulus criteria limits.
Gabriel test It is a pairwise comparison test using studentized maximum modulus. Unequal sample bounds. More powerful than Hochberg's test under unequal sample size. Liberal when sizes vary largely.
Waller-Duncan test Multiple comparison test based on t-statistic and uses Bayesian approach. Prior relative risk evaluation. Bayesian ratio allocation.
Games-Howell GH Based on the q-statistic distribution framework. When we have unequal variances and unequal sample sizes. Liberal test structure.
Dunnett's T3 and Dunnett's C Maintains strict control over the alpha significance level. Unequal variance and unequal sample size distributions. Strict protection filters.
Tamhane's T2 Conservative control adjustment framework. Unequal variance and unequal sample sizes. Conservative adjustment classification.
The Tukey-Kramer test Adjusted pairwise format analyzer. This is used when we have equal variances but unequal sample sizes. Unbalanced uniform variances.
The Miller-Winer test Equal distribution control matrices. When equal variances are assumed across metrics. Standard variance reference.
Multiple range (homogeneous subset) tests Based on q statistics formulas. Clustered range analytics. Subset variance grouping blocks.

Bonferroni-adjusted multiple t-tests (Dunn)

It is a simple type of multiple comparison tests. It keeps the family-wise error to a fixed value. Family-wise error is defined as the probability of having at least any one of the test results in a Type I error out of a series of tests.

Positives
  • It is preferred primarily when there are a low number of total comparisons under evaluation.
  • Highly versatile and suitable for both non-pairwise as well as traditional pairwise comparison paths.
Limitations
  • It does not have enough statistical power to effectively detect smaller meaningful variations as significant.

Sidak Test

The alpha significance level configuration computed for handling multiple comparisons runs noticeably closer than the traditional Bonferroni boundaries.

Limitations
  • It properly controls the family-wise error rate specifically when the analytical comparisons run strictly orthogonal to each other.

Computational Implementation inside R Environment

The following example executes pairwise t-tests modified by a global Bonferroni adjustment calculation template against sample workspace observations:

# Define treatment grouping factors group <- c(4,1,4,4,2,1,2,1,4,3,2,2,3,1,4,3,3,2,2,1,3,4,3,1,4,2,4,3,1,2,1,2,3,4,4,2,2,1,4,3,1,2,4,1,3,2,3,1,2,3,1,4,1,2,3,3,2,3,3,2,4,2,4,3,4,4,1,2,1,4,2,3,3,3,2,4,2,1,1,3,3) # Define dataset vector entries tracking event duration metrics survival_in_month <- c(145,198,122,108,61,82,121,191,191,109,82,74,85,193,65,173,129,107,157,175,173,169,104,2,129,181,189,9,190,70,7,190,86,141,64,134,121,200,177,59,26,137,64,18,55,68,30,147,30,151,86,83,22,169,103,13,2,183,27,60,132,109,172,13,44,124,87,31,130,90,69,37,57,195,4,195,62,144,35,187,193) # Compute mean outputs clustered by index labels tapply(survival_in_month, group, mean) # Execute baseline descriptive comparison filter matrixes pairwise.t.test(survival_in_month, group, p.adj = "bonferroni")

Academic Literature References

  1. Kucuk, U., Eyuboglu, M., Kucuk, H. O., & Degirmencioglu, G. (2016). Importance of using proper post hoc test with ANOVA. International Journal of Cardiology, 209, 346.
  2. Rice, W. R. (1989). Analyzing tables of statistical tests. Evolution, 43(1), 223-225.
  3. Wallenstein, S. Y., Zucker, C. L., & Fleiss, J. L. (1980). Some statistical methods useful in circulation research. Circulation Research, 47(1), 1-9.
  4. Nakagawa, Shinichi. "A farewell to Bonferroni: the problems of low statistical power and publication bias." Behavioral Ecology 15, no. 6 (2004): 1044-1045.
  5. Elliott, A. C., & Hynan, L. S. (2011). A SAS macro implementation of a multiple comparison post hoc test for a Kruskal Wallis analysis. Computer Methods and Programs in Biomedicine, 102(1), 75-80.
  6. Brown, A. M. (2005). A new software for carrying out one-way ANOVA post hoc tests. Computer Methods and Programs in Biomedicine, 79(1), 89-95.
  7. Blakesley, R. E., Mazumdar, S., Dew, M. A., Houck, P. R., Tang, G., Reynolds III, C. F., & Butters, M. A. (2009). Comparisons of methods for multiple hypothesis testing in neuropsychological research. Neuropsychology, 23(2), 255.
  8. Dinno, A., & Dinno, M. A. (2015). Package dunn.test.