The performance of robust procedures for multiple comparison tests under heteroscedasticity in psychological research

Loading...
Thumbnail Image

Authors

Liu, Linsey

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Multiple comparison test (MCT) strategies are widely used in psychological research following an omnibus F test (aka analysis of variance; ANOVA). Traditional methods such as Bonferroni, Tukey’s Honestly Significant Difference, and Dunnett’s test aim to control the familywise error rate (FWER) but rely on the assumption of homogeneity (i.e., variances of residual scores are the same across the levels of an independent variable), which is often violated in practice. To address these violations, robust alternatives—such as those incorporating sandwich estimators or the Plug-In procedure—have been proposed. However, due to differences (e.g., k setting, variance ratio (VR) setting, the sample size setting) in prior simulation settings, it remains unclear which procedures perform best under realistic conditions involving heteroscedasticity. This study systematically evaluated the robustness of MCT strategies, including classical MCT procedures (Bonferroni’s test, Tukey’s test and Dunnett’s test) and robust procedures (sandwich estimator and plug-in procedure) via Monte Carlo simulations under 51 manipulated conditions, including 12 balanced conditions with the same sample size (3 levels of group size * 3 levels of VR + 3 homogeneity conditions) as control group, 19 unbalanced conditions with 3 different group sizes (6 levels of combinations between group sizes & variances * 3 levels of VR + 1 condition where VR=1), 10 unbalanced conditions with 1 extremely small group (3 levels of combinations between group sizes & variances * 3 levels of VR + 1 condition where VR=1), and 10 unbalanced conditions with 1 extremely large group (3 levels of combinations between group sizes & variances * 3 levels of VR + 1 condition where VR=1). Both classical and robust MCT methods were examined across four key metrics: Type I error rate, confidence interval (CI) exclusion criterion, width of CI, and statistical power. Results showed that classical methods (Tukey, Dunnett, and Bonferroni) performed well in balanced conditions but exhibited an inflated Type I error rate and reduced power under variance heteroscedasticity or unequal sample sizes. In contrast, robust procedures maintained more stable power and better control of the Type I error rate across varied conditions. The CI results further revealed that robust methods provided more flexible and accurate adjustments for interval width associated with a better coverage rate of the true parameter value, particularly in complex and unbalanced designs. Among robust strategies, Tukey-HC2, Tukey-HC3, and Dunnett-PI consistently demonstrated the best trade-off between power and control of Type I error rate. Tukey combined with Heteroscedasticity-Consistent (HC) estimators minimized Type I error rate, while Dunnett paired with the PI procedure maximized power. Games-Howell effectively limited false positives but at the cost of lower power, making it more suitable when flexibility is prioritized. Overall, the findings underscore the importance of selecting MCT procedures that are both statistically powerful and robust to assumption violations. This study offers practical recommendations for psychological researchers, highlighting the advantages of robust methods in enhancing the accuracy and reliability of post hoc inference.

Description

Keywords

Multiple Comparison Tests, Robust Procedures, Sandwich Estimator, Plug-In Procedure, Heteroscedasticity, Monte-Carlo Simulation

Citation