Is it OK to dichotomize? A research dialogue
通过一场研究对话,探讨在实验研究中将连续预测变量二分化(如中位数分割)是否合理,分析了其统计功效损失和假阳性风险,并提出了在特定条件下可接受的观点。
Experimental consumer research—and experimental psychological research in general—often involves the study of measured predictor variables (e.g., need-for-cognition, materialism, impulsivity) in addition to conventional factors that are experimentally manipulated. Until fairly recently, it was common for experimental consumer researchers to dichotomize their measured predictor variables—typically through a median split—and to treat the dichotomized variables as discrete factors in standard ANOVAs. This practice was generally considered standard and acceptable. Around the early 2000s, however, a series of methodological articles, including some co-authored by researchers within our field (e.g., Irwin & McClelland, 2003; MacCallum, Zhang, Preacher, & Rucker, 2002), strongly criticized the dichotomization of continuous predictor variables. According to these critiques, not only is dichotomization causing a substantial reduction in statistical power (thus increasing the chance of type-II error), it can also create spurious results (i.e., type-I error) under certain conditions (see Maxwell & Delaney, 1993). Based on these critiques, a strongly worded “death to dichotomizing” was called for in an influential editorial appearing in the Journal of Consumer Research (Fitzsimons, 2008). This editorial and the critiques that it was based on have had strong effects on practices within our field. Today, rarely do we see continuous predictor variables analyzed with median splits combined with standard ANOVAs. Instead, multiple regressions combined with “spotlight” analyses (Aiken & West, 1991)—and most recently “floodlight” analyses (Spiller, Fitzsimons, Lynch, & McClelland, 2013)—have become de rigueur for the analysis and reporting of continuous predictor variables. It is against this backdrop that the present Research Dialogue is set. In their anchor article Dawn Iacobucci, Steven Posavac, Frank Kardes, Matthew Schneider, and Deidre Popovich (hereafter, IPKSP) challenge the current wisdom that dichotomization of continuous predictor variables is clearly be to avoided in experimental research. They summarize different reasons why researchers may still find dichotomization appealing. More importantly, they revisit the Maxwell and Delaney (1993) simulation result—which many critiques of dichotomization rely on—that had suggested that dichotomization can increase type-I error under certain conditions. IPKSP argue that the conditions that Maxwell and Delaney identified in their simulation were in fact quite unrealistic and atypical of actual behavioral research. In their article IPKSP offer their own simulations to evaluate the effects of dichotomizing a continuous predictor variable under a much broader range of conditions than Maxwell and Delaney (1993) originally studied. Their results suggest that while dichotomization does reduce statistical power, as is well established, it does not produce spurious results (type-I error) unless there is substantial multicollinearity among the predictor variables. Even if there is multicollinearity, the distortion in the observed estimates appears to be rather small. Based on these simulation results, IPKSP argue that the field should soften its stance against dichotomization, and allow researchers to dichotomize continuous predictors if two conditions are met: (a) the primary interest is in group differences rather than in individual differences; and (b) there is no multicollinearity among the predictors. According to IPKSP, the latter condition effectively covers one of the most common applications of dichotomization: the factorial crossing of a dichotomized continuous predictor with an experimentally manipulated variable. In response to IPKSP's piece, two commentaries were invited, both from authors who have been critical of dichotomization. In the first commentary Derek Rucker, Blakeley McShane, and Kristopher Preacher (RMP) point out that while IPKSP focus on their simulations' result that dichotomization does not increase the chance of type-I error when there is no multicollinearity, one should not forget their complementary finding that dichotomization does increase the chance of type-I error when there is multicollinearity. RMP additionally remind us that dichotomization does increase the chance of type-II error—a cost that they consider substantial and not to be ignored. Given this and other costs, RMP maintain that regression-based analyses of continuous predictors should remain the normative procedure. To help researchers follow this normative procedure, RMP provide a helpful tutorial on effective graphical representations of continuous data and an excellent primer on how to handle various statistical issues that may be encountered by researchers working with continuous predictors. In the second commentary, Gary McClelland, John Lynch, Julie Irwin, Stephen Spiller, and Gavan Fitzsimons (hereafter, MLISF) vehemently dispute IPKSP's suggestion that dichotomization is acceptable under certain conditions. After summarizing the basic statistical case against dichotomization, MLISF challenge the various nonstatistical arguments mentioned by IPKSP in support of the practice. They reject IPKSP's suggestion that the loss of statistical power induced by dichotomization effectively makes results more conservative. MLISF argue that with dichotomization, results are not necessarily more conservative. Results may in fact be less conservative in that sometimes regression slopes that are not statistically significant can still produce differences in means that appear significant only after dichotomization. They also explain that at a field-wide level—as opposed to the individual-paper level—conditions that lower statistical power on average increase the likelihood that a published body of knowledge contains false positive results. Moreover, MLISF have serious concerns about IPKSP's simulations, asserting that they were technically flawed with respect to how the effects of dichotomization on estimates of interactions were modeled. In addition, MLISF suggest that the results that IPKSP report, which are aggregated across the numerous conditions of their simulations, mask substantial distortions in the estimates that take place in some of the conditions—distortions that are offset in the aggregate by other distortions going in the opposite direction. Rejecting the IPKSP simulation results as a whole and offering their own derivations and calculations, MLISF conclude that dichotomization remains a “bad idea.” In their detailed rejoinder, IPKSP suggest that RMP's primary focus on the cost of dichotomization in terms of type-II error essentially amounts to a concession that the practice does not entail an increase in type-I error when the predictors are uncorrelated, which is IPKSP's main result. IPKSP argue that given that the field and the journals tend to be more concerned with type-I error (than with type-II error) for typical theory-testing papers, their simulations' primary result is noteworthy. They further maintain that there is nothing inherently wrong with dichotomizing continuous variables, and that researchers should generally be allowed to choose among their preferred analytical tools. IPKSP go on to respond to MLISF's criticisms point by point. They note that just like any other statistical method, multiple regression (with continuous predictors) raises its own set of issues and limitations. According to IPKSP, the issues of loss of statistical power are not as critical as MLISF suggest they are. IPKSP vigorously defend their original simulations and suggest that it is MLISF who are mistaken about how to best simulate the effects of dichotomization on estimates of interactions. With respect to their aggregate results masking significant distortions under certain conditions, IPKSP counter that their results do provide a full summary picture and that readers are free to regenerate and examine all results using the programming codes that are included with the original article. IPKSP conclude by reaffirming that the present methodological canon that dichotomization is always to be avoided is too extreme, and it needs to be reconsidered. The intensity of this exchange suggests that the field was indeed in need of a detailed discussion of the appropriateness of dichotomization in consumer research. Ultimately, I believe that this Research Dialogue offers our field a more refined understanding of the issues underlying this seemingly simple but evidently contentious subject. The discussion is probably not over yet.1