The Diamond Ratio (DR), Our Estimate of Heterogeneity: Now Published Online

I recently posted about our DR article being accepted by BJMSP. It has now been published online, here.

It’s behind a paywall, but here is the full pdf that we are allowed to share; it has only a few limitations, including no local file download.

Again, well done Max!

Enjoy,

Geoff

The Diamond Ratio (DR), Our Estimate of Heterogeneity: Accepted for Publication

I’m excited to report that Max Cairns’s PhD work on the Diamond Ratio (DR) has been accepted for publication by the British Journal of Mathematical and Statistical Psychology. The preprint of the final accepted version is here.

Our preprint, as accepted by BJMSP

The Original DR Blog Post

…was in 2018: Measuring Heterogeneity in Meta-Analysis: The Diamond Ratio (DR)

It includes a brief intro to the Fixed Effect (FE) and Random Effects (RE) models for meta-analysis, and to heterogeneity. When there’s heterogeneity, the diamond depicting the 95% CI on the result of an RE meta-analysis is likely to be longer than the FE diamond.

The DR is simply the ratio of those two diamond lengths. If DR = 1 there’s little or no heterogeneity; DR of, say, 1.5 suggests moderate heterogeneity. Larger DR, more heterogeneity.

Bob and I introduced the DR in Chapter 9 of ITNS. We wanted to discuss heterogeneity, while avoiding the complexities of conventional estimates of heterogeneity (Q, I2, Τ2). If both the RE and FE diamonds are displayed at the bottom of a forest plot, the DR can easily be eyeballed (see figure below).

A CI on the DR

Last year I posted about Max’s success in developing a very good approximate CI on the DR:

A Confidence Interval for the Diamond Ratio: Estimation of Heterogeneity in Meta-Analysis

The Paper Accepted for Publication

It’s easy to eyeball the DR—the ratio of the lengths of the two red diamonds in this figure from the preprint:

Forest plot summarising a meta-analysis performed on data in Figure 9.2 of ITNS. Both RE and FE diamonds are displayed, in red. The DR and its CI are reported (lower left) to be 1.40 [1.00, 3.09]. Eyeballing the ratio of the lengths of the diamonds agrees with that reported value of DR.

The preprint describes Max’s investigations of seven (!) approaches to calculating a CI for DR. These are all approximations, but, as the preprint reports, Max carried out extensive simulations to evaluate all of these. His main focus was on coverage. He found that the ‘Sub-Q’ approach has excellent coverage (very close to 95%) across a very wide range of situations.

Max’s short description of this CI is that it is the “Substitution CI with the Q-profile τ2 interval estimator”. No, that’s not totally clear to me either, but see the preprint for the full story, including the impressive (imho) range of simulation results. There are links to all the data and results, and to R software for calculating the CI on DR.

The red line below the RE diamond in the figure represents the length of the 95% prediction interval (PI) for the population effect sizes estimated by different individual studies. In other words, it indicates the likely extent of spread of these true effect sizes. It’s a further way to visualise the likely extent of heterogeneity. The length of the PI is 0.29, as reported in the figure.

esci in jamovi

Bob has included calculation of the CI on DR in the current beta of esci in jamovi, available here.

Visualisation is Understanding

Well, very often it’s a big help, and not only for beginning students. Visualisation is a focus all through esci and ITNS. We hope DR visualisation will help students and even researchers achieve a better understanding of heterogeneity.

As usual, it’s highly valuable to have a CI as well as a point estimate—in this case, of heterogeneity. Unless k, the number of studies in the meta-analysis, is quite large, the CI on DR is likely to be long, as it is in the figure above. That CI extends from the minimum, 1, to more than 3. With only k = 10 studies, all we can say is that heterogeneity is most likely between zero and very large. In other words, the amount of heterogeneity could be pretty much anything. That’s unfortunate, but it’s better to know this than be misled by any seemingly precise point estimate.

Well done Max! (And supervisor, Luke.)

Geoff

The Myth (?) of the Lucky Golf Ball Lives On, Alas

The APS has just given a kick along to what’s most likely a myth: The Lucky Golf Ball. Alas!

Golf.com recently ran a story titled ‘Lucky’ golf items might actually work, according to study. The story told of Tiger Woods sinking a very long putt to send the U.S. Open to a playoff. “That day, Tiger had two lucky charms in-play: His Tiger headcover, and his legendary red shirt.”

The story cited Damisch et al. (2010), published in Psychological Science, as evidence the lucky charms may have contributed to the miraculous putt success.

Laudably, the APS highlights public mentions of research published in its journals. It posted this summary of the Golf.com story, and included it (‘Our science in the news’) in the latest weekly email to members.

However, this was a misfire, because the Damisch results have failed to replicate, and the pattern of results has prompted criticism of the work. Read on…

The original Lucky Golf Ball study

Damisch et al. reported a study in which students in the experimental group were told—with some ceremony—that they were using the lucky golf ball; those in the control group were not. Mean performance was 6.4 of 10 putts holed for the experimental group, and 4.8 for controls—a remarkable difference of d = 0.81 [0.05, 1.60]. (See ITNS, p. 171.) Two further studies using different luck manipulations gave similar results.

The replications

Bob and colleague Tracy Caldwell (Calin-Jageman & Caldwell, 2014) carried out two large preregistered close replications of the Damisch golf ball study. Lysann Damisch kindly assisted them make the replications as similar as possible to the original study. Both replications found effects close to zero.

Meta-analysis of original and replication studies

Here’s Figure 9.8 from ITNS, a forest plot that shows results from six studies from the Damisch group (red), and the two Calin-Jageman & Caldwell replications (blue). The first of the red studies and both of the blue used the lucky golf ball task.

The clear difference between the overall red mean and overall blue mean is shown at the bottom on a Difference axis; it’s -0.77 [-1.15, -0.38], so a clear failure to replicate.

Not replicable, but citable

That’s the title of a 2018 post by Bob lamenting the common pattern of a striking original finding that continues to make waves, even while strong counter evidence from replications languishes in the shadows.

Below I’ve updated his figure showing citation counts for the original Damisch article and the Bob & Tracy replication article. The pattern has not improved these last three years!

The pattern of Damisch results

The six red CIs in the forest plot are astonishingly consistent, all with p values a little below .05. Greg Francis, in this 2016 post, summarised several analyses of the patterns of results in the original Damisch article. All, including p-curve analysis, provided evidence that the reported results most likely had been p-hacked or selected in some way.

Another failure to replicate

Dickhäuser et al. (2020) reported two high-powered preregistered replications of a different one of the original Damisch studies, in which participants solved anagrams. Both found effects close to zero.

All in all, there’s little evidence for the lucky golf ball. APS should skip any mention of the effect.

What next?

Open Science practices will help.  Perhaps high quality replication articles can be marked with big badges and trumpet fanfares? With everything online, it should be possible to add annotations to original articles and provide links to later replications, no doubt with original authors having a right of reply. Meanwhile, we need to keep up the skepticism and eternal vigilance.

Geoff

Calin-Jageman, R. J., & Caldwell, T. L. (2014). Replication of the Superstition and Performance Study by Damisch, Stoberock, and Mussweiler (2010). Social Psychology, 45(3), 239–245. https://doi/10.1027/1864-9335/a000190

Damisch, L., Stoberock, B., & Mussweiler, T. (2010). Keep Your Fingers Crossed! Psychological Science, 21(7), 1014–1020. https://doi/10.1177/0956797610372631

Dickhäuser, O., Heinze, A., Hamm, M. L., Bales, A. S., Bellmann, S. A., Böger, D., et al. (2020). Zwei teststarke, präregistrierte Replikationsstudien zum Einfluss von Glück auf kognitive Leistung (Two high-powered preregistered replication studies on effects of superstition on cognitive performance). Zeitschrift für Pädagogische Psychologie, 34, 51-60. https://doi.org/10.1024/1010-0652/a000263

What Should We Call Our Estimate of Cohen’s δ: d-unbiased, Hedges’ g, or Something Else?

In ITNS we used ‘dunbiased’ to refer to the debiased estimate of Cohen’s δ, which is Cohen’s standardised effect size in the population. In UTNS I used ‘dunb’. But now ‘Hedges’s g’ seems to be gaining currency as a label for that debiased estimate, despite g having been introduced by Larry Hedges back in the 1980s with a different meaning.

A bit of background

Cohen’s d for two independent groups, of size n1 and n2, with means M1 and M2 and SDs of s1 and s2 is

d = (M1M2) / (standardizer)

where ‘standardizer’ is some SD we choose as an appropriate unit of measurement for d. The numerator (difference between the means) is the effect size of research interest in original units and d is that ES re-expressed as a number of SDs; it’s a kind of z score.

Choice of standardizer is critical: d needs to be interpretable in the context. If our data are IQ scores on a well-established test, we might choose as standardizer σ = 15, the SD in the test’s reference population. But usually we’ll need to choose an estimate, calculated from the data, as standardizer. For two groups, it’s common to assume homogeneity of variance and use sp, the pooled estimate of the population SD. If one group is a control group, we might choose the SD of that group as standardizer, thus avoiding the assumption. Other choices are possible.

Unfortunately, d is a biased estimate: it overestimates δ, especially for small samples. A simple calculation debiases d. Of course, to interpret a value of d we need to know what standardizer was used and whether the reported value has been debiased. My question: What symbol should we use for debiased d?

A history of confusing labels

In 2011, in UTNS (p. 295), I wrote:

You’d think something as basic as dunb would have a well-established name and symbol, but it has neither. … In the early days the two independent groups d calculated using sp [pooled SD for the two groups, assuming homogeneity of variance] as standardizer was referred to as Hedges’ g.  For example the important book by Hedges and Olkin (1985), which is still often cited, used g in that way, and used d for what they explained as g adjusted to remove bias.  So their d is my dunb.  By contrast, leading scholars Borenstein et al. (2009) swapped the usage of d and g, so now their d is the version with bias, and Hedges’ g refers to my dunb.  Maybe hard to believe, but true.  The CMA [meta-analysis] software also uses g to refer to dunb.  In further contrast, Rosnow and Rosenthal (2009) is a recent example of other leading scholars explaining and using Hedges’ g with the traditional meaning of d standardized by sp and not adjusted to remove bias.  Yes, that’s all surprising, confusing, and unfortunate. 

Larry Hedges is one of the authors of Borenstein, et al. (2009), so presumably he supported the swapping of the labels. I asked him about these issues and he kindly replied with an account of the history. In his foundational articles of 1980-82 he used g for the biased estimate, to honour meta-analysis pioneer Gene Glass, and gU for the unbiased version. Then from around 1985 he started using d for the unbiased estimate, to correspond with δ (delta, Greek ‘d’). He reports that he doesn’t know who started using g for the unbiased estimate but that, by 2009, his co-authors felt that they should go with what seemed to have become standard practice.

Where to now?

My informal impression—I could be wrong—is that ‘Hedges’ g’ is increasingly being used for debiased d.

Bob and I need to decide what we’ll do in ITNS2 and in esci. Specifically, should we stick with ‘dunbiased’, or switch to using ‘Hedges’s g’. (Whichever we choose, we’ll no doubt note that both terms are in use.)

Despite the possible messiness of a long word as subscript, I’m currently leaning towards sticking with dunbiased. My thoughts:

  1. ‘Cohen’s d’, or simply ‘d’, is overwhelmingly the term used to denote the standardized ES. It’s used to introduce and explain the idea, and in journal articles—sometimes even if debiased values are reported. Further, dunbiased signals a particular variant of d, and even explains its key property—being an unbiased estimate. Guessing would probably give reasonable understanding.
  2. I suspect most researchers have heard of d, have interpreted values and perhaps used d in their own research, even if they don’t know about debiasing—which anyway isn’t an issue for N more than, say, 50. Many fewer would have heard of Hedges’ g, or be able to link it to d, let alone say how it relates to d. Both those links would need to be explained, and taught. The change of letter symbol seems arbitrary; there is no way to guess.
  3. It’s common (and useful) to refer to ‘the d family’ of standardized effect size measures. How strange that the most commonly needed member of that family is labeled ‘g’.
  4. It’s a great convention that a Roman letter estimates the corresponding Greek letter. So M, s, and r estimate µ, σ, and ρ respectively. Therefore it’s great that δ is widely used for the population value of Cohen’s d. Using g for the sample value suggests we’re estimating γ, which is never used. How weird to have to explain that the best estimate of δ is g.
  5. A mathematical statistician would sidestep all the above by using “delta hat” for the estimate, but I don’t think that’s a good universal solution for psychology, or many other research fields.
  6. In medicine, SMD, for “standardised mean difference” is widely used and a reasonable acronym. However, it can refer to a population value or sample estimate, and very often we’re left to wonder whether bias has been removed.
  7. On the other hand, it’s useful for a subscript to signal how d is calculated, perhaps ds when sp is the standardizer and we assume homogeneity of variance, and dC when the SD of the Control group is standardizer and we avoid that assumption. Using d and g permits subscripts to tell us about the standardizer. However, I don’t think any strong conventions have emerged as to which subscripts tell us what.

Given all that, I’m currently preferring dunbiased. However, has g become unstoppable? If so, the complexity of d, g, and δ is just one more baffling inconsistency we have to explain to bemused students.

Please let me have your thoughts.

Geoff

Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2009). Introduction to meta-analysis. Chichester, UK: Wiley.

Cohen, J. (1969). Statistical power analysis for the behavioral sciences. New York: Academic Press.

Hedges, L. V., & Olkin, I. (1985). Statistical methods for meta-analysis. Orlando, FA: Academic Press.

Which Standardised Effect Size Measure Is Best When Variances Are Unequal?

A great new preprint by Marie Delacre (at Université Libre de Bruxelles, marie.delacre@ulb.be) and colleagues (Daniel Lakens, Christophe Ley, Limin Liu, & Christophe Leys) throws valuable light on this question.

The title is: Why Hedges’ gs* based on the non-pooled standard deviation should be reported with Welch’s t-test

The issue is important for Bob and me as we work on ITNS2 and esci in jamovi, so I was an avid reader. I sent comments and questions and have had a quick and generously detailed response from Marie. She intends to revise the paper around September. I suspect she would be happy to have further comments.

Below is my take on the preprint. In brief, the authors report numerous simulations to investigate the properties of 8 (!) standardised ES measures, focussing on unequal variances and departures from normality.

When variances are equal: Two familiar ES estimates

With two independent groups and assuming homogeneity of variance we usually use Cohen’s d, being the difference between sample means divided by the pooled SD, sp. The pooled SD is the standardiser, the unit of measurement for d. Then a simple adjustment gives us dunbiased, also called Hedges’ g, as an unbiased estimate of δ, the population effect size (ES). Cohen’s d and Hedges’ g are the first of the ES measures investigated.

(In the preprint, the 8 ES measures are indicated as Cohen’s ds and Hedges’ gs, etc, with ‘s’ subscript. These seem redundant and I understand may be removed in the revised version.)

When variances are not equal: Six further ES estimates

Sometimes it’s unjustified, or questionable, to assume population variances are equal. For example, a treatment often increases the variance as well as the mean, compared with the Control condition. It may then make sense to use sC, the SD of the Control group, as standardiser, to get Glass’s d, which becomes Glass’s g when debiased. These are the third and fourth of the ESs studied.

When variances are unequal, we use Welch’s t test:

The denominator is an estimate that weights the two sample variances by sample size, with the larger group receiving the smaller weight. For inference, as with a t test, that’s correct—think of the formula for the SE.

Shieh (2013) proposed using a standardised ES measure based on a standardiser closely related to the denominator in the equation for t‘. In a comment (Cumming, 2013) I argued that Shieh’s d was pretty much uninterpretable: Among other problems, it didn’t estimate an ES in any existing population, and its value was greatly dependent merely on the relative sizes of the two samples. I recommended against using it.

Delacre and colleagues cited my comment, but did include Shieh’s d and Shieh’s g (the unbiased version) for completeness and in line with earlier work of theirs on inference (e.g. Delacre et al., 2017) that advocated use of Welch’s t.

However, inference should not dictate choice of standardiser: We sometimes need a standardiser not based on the SE appropriate for inference, e.g. in the simple paired design, as discussed in ITNS, pp. 207-208.

Finally, consider

which bases the standardiser on the average of the two sample variances, whatever the sample sizes. Again, it’s challenging to interpret because it doesn’t estimate a population ES for any existing population, but at least it’s not dependent on relative sample sizes. Cohen’s d* and its unbiased version, Hedges’ g*, complete the 8 ES estimates investigated by Delacre and colleagues.

Results and recommendations

The simulations explored bias and variance of the 8 ES measures for a range of pairs of population variances, pairs of sample sizes, and normal and 3 distinctly non-normal population distributions: a massive project giving a rich trove of information about the robustness of 8 measures. There are numerous tables and figures of estimates of bias and variance to pore over.

The authors’ conclusions:

  • “Because the assumption of equal variances… is rarely realistic… both Cohen’s d and Hedges’ g should be abandoned.” (p. 10)

That’s arguable. I’m not convinced the assumption is rarely realistic. (It’s also very often made, even if sometimes it shouldn’t be.) The emphasis should be on informed judgment in context rather than simply abandoning these two most familiar estimates. In addition, when population variances are equal, Hedges’ g performs very well. It’s also familiar and readily interpretable.

  • Shieh’s d and Shieh’s g generally perform poorly and are not recommended.

That’s a relief and what I expected. Let’s not consider them further.

  • “We do not recommend using [Glass’s d or g].” (p. 28)

I suggest that Glass vs something else is the choice that most clearly should be based on the context. Does it make sense to use the SD of one group, often the Control group, as the standardiser? If so, we should do so, unless there are very strong reasons against. We should use choice of sample sizes and perhaps other strategies (transform the DV to reduce departure from normality?) to minimise any disadvantage of the Glass’s g estimate. The simulation results give valuable guidance on when we might be concerned and what strategies might help.

  • “The measure … we believe performs best across scenarios is Hedges’ g*.” (p. 28).

This conclusion is expressed in the preprint’s title: Why Hedges’ gs* based on the non-pooled standard deviation should be reported with Welch’s t-test. The authors draw this conclusion despite having noted the wide criticism of Cohen’s d* (and by implication Hedges’ g*) because the standardiser is not the SD of an existing relevant population, so may be difficult to interpret.

Interpretability as the primary requirement for a standardised ES

When should we transform from an original to a standardised measure? What’s the purpose? As the authors note (pp. 3-4), a standardised measure can assist (i) interpretation of results in context and (ii) comparison of results for DVs with different original measures, for example using meta-analysis. It’s also (iii) useful when planning studies, whether using precision for planning or statistical power.

Above all, I’d argue, we need to be able to make sense of any point estimate—what is it estimating, what’s the unit of measurement, what does its magnitude tell us in the context? We also need an interval estimate to tell us the precision.

Hence my above comments that I would consider Hedges’ g and perhaps Glass’s g first, and contemplate Hedges’ g* only if those first two seemed seriously problematic and I couldn’t find a way to make them acceptable in context.

Estimation and assessing robustness

I’m looking for quantitative guidance about the likely bias in the point estimate, and error in the CI length of for example my favourite, Hedges’ g, in some context. If bias is likely to be 1-2% or a nominally 95% CI to have 92% or 96% coverage in the context, then I may stick with Hedges’ g. I’d have in mind the dance of the means and dance of the CIs: Replicate and most likely get a quite different point and interval estimate, so let’s not fuss too much about tiny biases. Within limits!

The Delacre simulations explore an admirably wide but realistic range of differences in sample sizes and variances, and departures from normality that are fairly extreme. I suspect the authors’ main strong conclusion in favour of Hedges’ g* is driven largely by big bias and variance problems found with the more extreme cases, although I’m not sure the extent that’s true.

However, if I’m dealing with g values less than 1 or 1.5, as often in psychology, and the sample sizes are within a factor of 2, how large is the likely bias? How close to 95% is the likely coverage of CIs? The robustness results are gold, and can answer many such questions, but will be most useful when re-expressed with such questions in mind. Further analysis and perhaps further simulations may be needed to give a full picture in terms of CI lengths and coverages. Then we’d have a wonderfully usable and valuable resource.

The title

Currently this is Why Hedges’ gs* based on the non-pooled standard deviation should be reported with Welch’s t-test’. If we want a p value, then Delacre et al. (2017) make a strong case for routinely preferring Welch’s t test over the conventional t test that requires homogeneity of variance: Little to lose if variances are equal and much to gain if not.

However, choice of standardised ES measure is a quite different question. Also, the formula for Welch’s t (formula above for t‘) bears no relation to that for Hedges’ gs*, so I see no reason to link the two in the title, especially since Welch’s t test is scarcely considered in the preprint.

My preference would be to use the title of this blog post, or something like: Cohen’s d and related effect size estimators: Interpretability, bias, precision, and robustness.

Finally

Marie Delacre has kindly indicated that she’s open to discussion as she and colleagues work on revisions. There may be future projects, perhaps focussed on CIs. Please add comments below, or send to her (marie.delacre@ulb.be) or me. Thanks.

Geoff

Cumming, G. (2013). Cohen’s d needs to be readily interpretable: Comment on Shieh (2013). Behavior Research Methods, 45, 968–971. https://doi.org/10.3758/s13428-013-0392-4

Delacre, M., Lakens, D., & Leys, C. (2017). Why psychologists should by default use Welch’s t-test instead of Student’s t-test. International Review of Social Psychology, 30 (1), 92–101. https://doi.org/10.5334/irsp.82

Shieh, G. (2013). Confidence intervals and sample size calculations for the standardized mean difference effect size between two normal populations under heteroscedasticity. Behavior Research Methods, 45, 955–967. https://doi.org/10.3758/s13428-012-0228-7

A Confidence Interval for the Diamond Ratio: Estimation of Heterogeneity in Meta-Analysis

A while back I posted (here) about the Diamond Ratio (DR), which is our simple visual indicator of the extent of heterogeneity in meta-analysis. (See ITNS, Chapter 9 for more on the DR.) I reported that Max Cairns, a PhD student in statistics at La Trobe University, and his supervisor Luke Prendergast were working on finding a confidence interval (CI) for the DR.

I’m delighted to report that they have now posted a preprint of their results here. We’d love to have your comments and suggestions.

Max explored six approaches to calculating a CI for the DR. He used simulation to investigate their properties, especially coverage, and identified two that give excellent CIs. He provides (here) R code to allow any researcher to calculate the CI on the DR for their own data, for a range of measures. All Max’s simulation materials are available on OSF here, so anyone can recreate or extend Max’s work.

Below is Figure 1 from the preprint, as an example of how the DR and its CI may be reported in a forest plot.

Figure 1. Forest plot summarising a meta-analysis performed on data in Figure 9.2 of ITNS. Eyeballing the ratio of the lengths of the two diamonds in the figure agrees with the reported value of DR = 1.399, which is an estimate of the amount of heterogeneity. The red line below the RE diamond represents the length of the associated prediction interval (PI) which, as reported in the figure, is 0.285.

In the figure, DR = 1.40 is reported along with three conventional measures of heterogeneity, all with CIs. Both the RE (Random Effects) and FE (Fixed Effect) diamonds are shown in the forest plot, so it’s easy to eyeball DR, which is simply the length of the RE diamond divided by that of the FE diamond. DR = 1 suggests little or no heterogeneity, and increasing values of DR suggest increasing heterogeneity. One vital message is given by the CI on the DR, which is [0, 3.09], so this meta-analysis, which integrates only 10 studies, can give us only a very imprecise estimate of heterogeneity.

Along with the DR, the figure reports the 95% prediction interval (PI) for true effect sizes as a further estimate of heterogeneity. Borenstein et al. (2017) advocated use of the PI, which is reported here to be 0.285. The red line segment just under the RE diamond pictures that length. Informally, that segment illustrates the likely extent of spread of true effect sizes. The PI is 4 x T, where T is the estimated population SD of true effect sizes. The very long CI reported for T indicates once again a very imprecise estimate of heterogeneity.

In the preprint we conclude that the DR, and its CI, can be valuable for students as they learn about meta-analysis, and for researchers as they interpret and communicate their meta-analyses.

It would be great to have any comments about Max’s work and the preprint. Thanks!

Geoff

Max: mrcairns994@gmail.com Geoff: g.cumming@latrobe.edu.au

Borenstein, M., Higgins, J. P., Hedges, L. V., & Rothstein, H. R. (2017). Basics of meta-analysis: I2 is not an absolute measure of heterogeneity. Research Synthesis Methods, 8, 5-18. doi:10.1002/jrsm.1230