The Diamond Ratio (DR), Our Estimate of Heterogeneity: Accepted for Publication

I’m excited to report that Max Cairns’s PhD work on the Diamond Ratio (DR) has been accepted for publication by the British Journal of Mathematical and Statistical Psychology. The preprint of the final accepted version is here.

Our preprint, as accepted by BJMSP

The Original DR Blog Post

…was in 2018: Measuring Heterogeneity in Meta-Analysis: The Diamond Ratio (DR)

It includes a brief intro to the Fixed Effect (FE) and Random Effects (RE) models for meta-analysis, and to heterogeneity. When there’s heterogeneity, the diamond depicting the 95% CI on the result of an RE meta-analysis is likely to be longer than the FE diamond.

The DR is simply the ratio of those two diamond lengths. If DR = 1 there’s little or no heterogeneity; DR of, say, 1.5 suggests moderate heterogeneity. Larger DR, more heterogeneity.

Bob and I introduced the DR in Chapter 9 of ITNS. We wanted to discuss heterogeneity, while avoiding the complexities of conventional estimates of heterogeneity (Q, I2, Τ2). If both the RE and FE diamonds are displayed at the bottom of a forest plot, the DR can easily be eyeballed (see figure below).

A CI on the DR

Last year I posted about Max’s success in developing a very good approximate CI on the DR:

A Confidence Interval for the Diamond Ratio: Estimation of Heterogeneity in Meta-Analysis

The Paper Accepted for Publication

It’s easy to eyeball the DR—the ratio of the lengths of the two red diamonds in this figure from the preprint:

Forest plot summarising a meta-analysis performed on data in Figure 9.2 of ITNS. Both RE and FE diamonds are displayed, in red. The DR and its CI are reported (lower left) to be 1.40 [1.00, 3.09]. Eyeballing the ratio of the lengths of the diamonds agrees with that reported value of DR.

The preprint describes Max’s investigations of seven (!) approaches to calculating a CI for DR. These are all approximations, but, as the preprint reports, Max carried out extensive simulations to evaluate all of these. His main focus was on coverage. He found that the ‘Sub-Q’ approach has excellent coverage (very close to 95%) across a very wide range of situations.

Max’s short description of this CI is that it is the “Substitution CI with the Q-profile τ2 interval estimator”. No, that’s not totally clear to me either, but see the preprint for the full story, including the impressive (imho) range of simulation results. There are links to all the data and results, and to R software for calculating the CI on DR.

The red line below the RE diamond in the figure represents the length of the 95% prediction interval (PI) for the population effect sizes estimated by different individual studies. In other words, it indicates the likely extent of spread of these true effect sizes. It’s a further way to visualise the likely extent of heterogeneity. The length of the PI is 0.29, as reported in the figure.

esci in jamovi

Bob has included calculation of the CI on DR in the current beta of esci in jamovi, available here.

Visualisation is Understanding

Well, very often it’s a big help, and not only for beginning students. Visualisation is a focus all through esci and ITNS. We hope DR visualisation will help students and even researchers achieve a better understanding of heterogeneity.

As usual, it’s highly valuable to have a CI as well as a point estimate—in this case, of heterogeneity. Unless k, the number of studies in the meta-analysis, is quite large, the CI on DR is likely to be long, as it is in the figure above. That CI extends from the minimum, 1, to more than 3. With only k = 10 studies, all we can say is that heterogeneity is most likely between zero and very large. In other words, the amount of heterogeneity could be pretty much anything. That’s unfortunate, but it’s better to know this than be misled by any seemingly precise point estimate.

Well done Max! (And supervisor, Luke.)

Geoff

The New APA Style: Try to Contain Your Excitement—and Watch Out for Dud Copies

This is a post for the nerds, fine people that we are. Actually, everyone needs to think about reporting style, especially for statistical stuff.

Seventh edition, 2020
Sixth edition, 2010

The APA Publication Manual is the bible of APA Style, also perhaps the bane of some students’ lives. In ITNS we used the 6th edition (APA, 2010). Now there’s a 7th edition (APA, 2020). I guess we need to switch for ITNS, second edition, tho’ I’m finding it hard to get enthusiastic.

Figures: captions out, headings in

For me, the biggest change is the demise of detailed figure captions placed below the figure. Now there’s a brief heading above a figure, and Note below to give the description. That matches style for tables, but I’m grieving already.

The OLD style for figures, as in ITNS

Figure 1.1. Support for Proposition A, in percent, as reported by the poll. The dot marks the point estimate, and the two lines display the margin of error (2%) either side of the dot. The full interval, from 51% to 55%, is the 95% confidence interval.

The NEW style for figures, as in ITNS2?

Figure 1.1

Support for Proposition A, as Reported by the Poll

Note. Support is expressed in percent. The dot marks the point estimate, and the two lines display the margin of error (2%) either side of the dot. The full interval, from 51% to 55%, is the 95% confidence interval.

What else is new in the 7th edition

A guide to the changes from 6th to 7th editions is here. A video that illustrates the main changes is here. It shows how the new style is a bit simpler—hooray—despite the new edition being more detailed and having more than 50% more pages than the old.

There are small changes to reference formats and how references are cited in text. Reference formats, and examples, are provided for more than 100 types of items, including social media items, blog posts and comments on blog posts, TED talks, songs, and just about every other type of item you can think of.

For the first time, format for student papers is discussed, with great scope given for instructors to specify how student work should be presented. Good.

JARS, the Journal Article Reporting Standards, are highly detailed, and now comprise a whole chapter (Chapter 3). As well as quantitative, they cover qualitative and mixed methods research. There’s particular mention of meta-analyses, and replication studies. Yay.

Only the super-nerds need notice that now only one space is required at the end of a sentence, not two. A massive saving of virtual trees?

Open Science?

There’s brief mention of Open Science badges, but the index includes no other item for ‘Open Science’. There’s no index entry for ‘preregistration’, but the ‘registration’ entry points to a couple of mentions of study registration, in particular of clinical trials. There’s advice on how to include access information about open data and open materials provided online. Overall, however, there is no strong advocacy of Open Science practices. Sadly.

The new Publication Manual: Beware counterfeits

Last year I bought a copy of the 2020 Manual on the Australian site of Amazon. I noticed small errors, for example incorrect italics. I enquired of APA, sent them photos of a few pages, and was told I had a counterfeit copy. It came from a third-party seller via Amazon. I returned the copy to Amazon, as requested, and gave the seller a blistering review—they even had the cheek to ask me to withdraw this as spoiling their business. Happily, APA sent me a replacement copy, although I suspect that’s not usual practice.

There’s now a blog post warning of counterfeits, as well as advice here and here on spotting duds.

Are journals using the new style?

Not surprisingly, APA journals are using the new style. The change is in progress. Skim through recent issues and see a mixture of articles using the old style for figures, and those using the new. For APS journals, the online guidelines for authors still refer to the old 2010 style. I enquired of Elaine Walker, Chair of the APS Publications Committee, who kindly explained that APS intends to swap to the 7th edition, and is currently planning the move. So, I guess Bob and I need to switch. Sigh. I’m already nostalgic for fully explanatory captions under figures.

Why we need reporting standards

Years ago some students and I scanned articles in economics, business, and chemistry journals to see how these disciplines reported statistical significance testing. This was pilot testing, never written up. We found that reporting was often sketchy and obscure. We saw text such as: “Figure 1 illustrates that only groups Groups A and C showed an effect.” It took investigation and guesswork to figure out that the authors had probably used a criterion of 2xSE to identify effects as existing or not. Dichotomous decision making, without even coming clean, stating the criterion being used, and admitting that any difference just less than 2xSE was regarded as not existing.

In Psychology, long-standing APA style requirements have meant that NHST is almost always more full reported. We should be told at least the summary statistics, what’s being tested, the df, the p value, and then a conclusion.

Now that researchers are using computer scanning of very large numbers of journal articles to study, for example, the use of NHST and the distribution of reported p values, it’s more vital than ever that authors report fully and comply with style standards in reporting statistical analyses.

At least to some extent, we all need to embrace our inner nerd. The 2020 version of course.

Geoff

jamovi—and esci—Just Keep Powering Ahead

Two beautiful numbers:

55,000 – the approx. number of times jamovi was downloaded in March

2,500 – the approx. number of times esci was added to jamovi in March

Each of these is about double the number for three months earlier! At this rate, everyone on Earth will have their own copy within a year or two—roughly speaking 😉

Of the 38 modules available in the jamovi library, esci is currently the 5th most popular—demonstrating that it’s fully usable despite still being in development. Hats off to Bob!

In case it’s new to you, jamovi is the free, open-source stats software that crushes SPSS. It’s even better with added esci—which is designed to go with the second edition of ITNS, currently in preparation.

To get started, see this post.

Even simpler than downloading jamovi—tho’ this is quick and easy—just click on the big green button at the jamovi home page to open jamovi directly in any browser. Then play. (The online version is experimental, and modules can’t yet be added.)

Enjoy, and please let’s have your comments and suggestions,

Geoff

What N Will Give Me the Precision I Want? Gordon’s New Pictures Tell All

We’re delighted to release precision for planning (PfP), the sixth component of esci web. This completes esci web as currently planned. To access, click esci web and then the precision for planning button. Please let’s know how you like it, and give us your suggestions.

Precision for Planning tells us what N we need to achieve the precision we’d like. It’s a much better way to plan than the traditional use of statistical power, which works only within an NHST framework. Far better to adopt an estimation framework (the new statistics) and use PfP.

For an intro to PfP, see Chapter 10 in ITNS. For more detail, see Chapter 13 in UTNS.

Two Independent Groups

For a two independent groups study, with two groups of size N, below is the PfP picture. Recall that MoE is the margin of error, which is half the length of a CI. I’ve set the slider at the bottom to target MoE = 0.50, meaning that I want to estimate the difference between the group means with a 95% CI having MoE of 0.50. In other words, each arm of the CI should be 0.50 in length.

precision for planning from esci web. The large slider marks our selected target MoE of 0.50. Two groups each with N = 32 will, on average, give a 95% CI with MoE no more than 0.50. The lower curve is the distribution of MoE lengths when N = 32.

The lower axis is marked in units of population SD, which we can think of as units of Cohen’s d. The cursor marks a target MoE of 0.50 in those units.

The black curve shows how required N increases dramatically as we aim for smaller values of MoE–in other words, greater precision and a shorter CI. Use this curve to investigate how N trades with likely precision.

The small curve at the bottom shows how MoE varies for N = 32. It’s usually close to 0.50, but can be as short as 0.40 or long as 0.60, and occasionally even a little outside that range. Use the large slider to move the cursor and see the MoE distribution for other values of target MoE and N.

The figure gives us a handy benchmark, worth remembering: Any study with two independent groups of size 32 will estimate the difference between the group means with a 95% CI that has MoE of 0.50, on average.

With 99% Assurance

The black curve can only give us N for MoE that’s sufficiently small on average. But we can do better. The red curve, below, tells us the N we need to achieve target MoE with assurance of 99%. This is the N that gives MoE smaller than target MoE on at least 99% of occasions. The grey curve reminds us of the ‘on average’ curve–the black curve in the figure above.

precision for planning from esci web. The red curve shows N required with 99% assurance, meaning that MoE will be no more than target MoE on at least 99% of occasions. Two groups each with N = 44 will give MoE no more than 0.50 on at least 99% of occasions. The small lower curve is the distribution of MoE lengths when N = 44. Only the tiny right tail, area at most 1%, exceeds 0.50.

The Paired Design

precision for planning supports PfP for what are probably the two most common designs: two independent groups, and the paired design. The paired design, with a single repeated measure (for example Pretest-Posttest) has the advantage, where it is possible and appropriate, of usually giving higher precision. The critical feature is the correlation in the population between the two measures, such as Pretest and Posttest. Higher correlation gives a shorter CI on the paired difference and therefore higher precision.

To use PfP we need to specify a value for ρ (Greek rho), the population correlation. Ideally, previous research gives us a reasonable estimate we can use; otherwise we might have to guess. For research with human participants, typical values are often around .6 to .9.

Here’s a PfP picture for the Paired Design, with ρ set to .70.

precision for planning for the Paired Design, with ρ set to .70. Target MoE is 0.50. The grey curve tells us a single group of N = 12 will give MoE no more that 0.50, on average. The red curve shows that N = 21 will achieve that target MoE with assurance. The lower curve is the distribution of MoE lengths when N = 21.

The red curve shows us that a single group of N = 21 suffices for target MoE = 0.50 with assurance, when ρ = .70. Compare with two groups of N = 44 for the independent groups design. Great news!

However, as you might guess, N is highly sensitive to ρ. For ρ = .60 we need N = 25, but for ρ = .80 we need only N = 16 (or N = 9, on average).

It’s wonderful that precision for planning makes it easy to explore how N, target MoE, choice of design, and–for the Paired Design–ρ, all co-vary. Be fully informed before you choose a design and N!

Full esci web

Go to esci web and see all six components as here:

Main Menu page of esci web. Click at left to go to any of the six components.

Search the blog for ‘Gordon‘ to find three posts introducing the previous five components.

Please explore any and all of the six components. Send your bouquets to Gordon Moore. Your comments and suggestions to any of us.

Enjoy!

Geoff

Gordon Does It Again: See the Correlations Dance

Here comes one more goodie from Gordon Mooredance r. This follows his wonderful dances, introduced here, and three other goodies introduced here. Have a play with dance r, the latest component in esci web. Please tell us what you think.

From dance r. Correlation in the population (grey cloud at top) is .80. The scatterplot shows also the latest sample (blue dots), size N = 30, with correlation r = .83. That r is the top value (green dot with 95% CI) in the dance picture below. Earlier sample correlations have danced down the screen and collected into the r heap. CIs that don’t capture population correlation ρ (blue line) are red.

As you may recall, ITNS2 will be accompanied by Bob’s data analysis software, esci, in R, and Gordon’s web-based simulations and tools, all of which are based on, and go beyond, my Excel-based ESCI. Together the web-based goodies, now including dance r, comprise esci web, which you can open in your browser here. (Or use the ESCI menu above and choose esci web from the dropdown.) From today, esci web has five components, with one more to come.

dance r takes random samples from a bivariate normal distribution with chosen correlation ρ. As in the figure above, see the scatterplot of each sample, and also the r values from successive samples dancing down the screen. That’s the sampling variability of the correlation–and often it’s scarily large!

Playing yourself is *way* better than seeing the pic. A few things to try:

  • Watch the population cloud change for different ρ values
  • Explore the changing length and asymmetry of CIs for different r values
  • Watch the sampling distribution of correlations (the r heap) build
  • See how its skew changes with ρ
  • Investigate the capture percentage of 95% CIs
  • Study what changes, and how fast, as you change N

A key challenge for students–and researchers–is to build good intuitions about the extent of uncertainty, including the extent of sampling variability. dance r is a great arena in which to build those intuitions about correlation.

As I say, we’d love to have your feedback.

Enjoy.

Geoff

Goodies from Gordon: ‘distributions’, ‘d picture’, ‘correlation’–all part of ‘esci web’

I’m delighted to say we’re releasing three new goodies from Gordon Moore: distributions, d picture, and correlation. These follow his wonderful dances introduced here.

As I explained, ITNS2 will be accompanied by Bob’s data analysis software, esci, in R, and Gordon’s web-based simulations and tools, all of which are based on, and go beyond, my Excel-based ESCI. Together the web-based goodies comprise esci web, which you can open in your browser here. (Or use the ESCI menu above and choose esci web from the dropdown.) From today, esci web has four components, with perhaps two yet to come.

distributions, d picture, and correlation are visual statistical tools, developed in JavaScript. We’d love to have your feedback.

distributions: Explore normal and t distributions

See the curves, explore z scores, find areas, find critical values.

Normal distribution, z scores below, IQ scores (or whatever you choose) above.
Normal, and t distribution (df = 9). Watch the shape difference as df zooms up or down.

d picture: Explore Cohen’s d values visually

What does d = 0.2 look like? How much overlap of distributions? What about d = 0.5, 1.0, 1.5, …?

Cohen’s d = 0.5, with the area under E above mean of C shaded.

correlation: See scatterplots and eyeball r values

What do you think is the r value in each of these scatterplots?

——— Don’t read on just yet. Have an eyeball of the scatterplots. What is each r?

——— Last chance… look back up…

OK, the correlation is .3 in all cases. True, if possibly strange. (All the data sets come from a bivariate normal distribution, and in all cases the data set correlation is .3.)

Pro tip: Eyeball, or turn on, a cross through the means, as in lower right. Then eyeball the approximate comparative number of dots in (top right + lower left) quadrants and the (top left + lower right) quadrants. Correlation is a tussle between those first two (the matched quadrants) and the second two (the unmatched).

Investigate that and other cool things in correlation.

As I say, access esci web here, and please let us have your comments.

Enjoy,

Geoff

Gordon’s ‘dances’: Vivid Simulations Bring Statistical Ideas Alive

Bob and I are delighted to welcome Gordon Moore who joins us in working on the second edition of ITNS. Gordon, an independent tutor in computing, statistics and mathematics, is based in England, so our ITNS2 team of three now spans three continents.

We are now releasing Gordon’s dances in beta, and seek your feedback. Developed in JavaScript, dances opens in your browser via this link. ITNS2 will be accompanied by Bob’s data analysis software, esci, in R, and Gordon’s web-based simulations, all of which are based on, and go beyond, my Excel-based ESCI. The first and most important of Gordon’s simulations is dances, which replaces and goes beyond CIjumping in ESCI.

Below are four examples of dances bringing key statistical ideas alive. These are frozen images: It’s way more convincing watching the simulations dancing down the screen.

Getting started with dances:

  • Open dances in a browser
  • Click on the ‘?’ at top right in the control panel (left side of screen) to turn on popout tips, which give brief explanations when the mouse hovers over labels or controls.
  • Use the three big buttons. Play as you wish. Click ‘Clear’ to start again.

1. Variability: Very often larger than we think. Dance of the means.

Take repeated samples of size N = 20 from the pictured normally distributed population. Watch the pattern of values (blue open circles) jump around from sample to sample. Watch the means (green dots) from successive samples dance down the screen: So much variation, even with samples of size 20! This is the dance of the means.

2. Randomness is lumpy but, in the long run, totally predictable. Dance of the confidence intervals.

Place 95% CIs on each of the dancing means, again with samples of N = 20. CIs that don’t capture the population mean, mu (blue line), are red. In the short term, red CIs seem to come very haphazardly, sometimes rarely, sometimes in clumps. In the long term, however, very very close to 95.0%  of CIs will capture mu and 5.0% will be red.

This happens when CIs are all the same length, being based on the population SD, sigma, assumed known. Remarkably, it also happens when, as in the picture below, CIs vary in length because they are based on sample SDs, when sigma is assumed not known. Either way, we are seeing the dance of the CIs.

The falling means pile up to form the mean heap; means in the heap keep their colour, red or green. In the long run, the mean heap shape will closely match the theoretically expected, normally distributed, sampling distribution curve.

3. The Central Limit Theorem: Surprisingly close, even with tiny samples.

The central limit theorem states that, almost whatever the shape of the population distribution, the sampling distribution of sample means will be approximately normal. Furthermore, the larger the samples, the closer the sampling distribution will be to normal.

In dances you can draw whatever weird shape of population distribution you choose, then take samples of some chosen size, N, and compare the mean heap with the normal curve.

The figure below shows that, even with my hand-drawn, highly skewed population, and samples as tiny as N = 3, the mean heap is much less skewed than the population, and surprisingly close in shape to the symmetric normal curve.

4. The p value varies so widely it can’t be trusted: Dance of the p values.

Run a replication, exactly the same as the original experiment but with a new sample, and find that the p value is likely to be very different. The sampling variability of the p value is surprisingly large: Alas, we simply shouldn’t trust any p value.

The figure below shows the dance of the CIs and the corresponding p values—which vary from <.001 to more than .8! Deep blue patches mark p>.10, through to bright red patches for p<.001. This is the dance of the p values!

Population mean, mu, is 60, and SD, sigma, is 20. The null hypothesis is H0: mu0 = 50, so the effect size in the population is half of sigma, or Cohen’s delta = 0.50, conventionally considered to be a medium-sized effect. With N = 16, the power is about .50, which is typical for many research fields in psychology and some other disciplines.

The running simulation is way more vivid than any picture, especially when sounds are turned on, ranging from a bright trumpet for p<.001 down to a deep trombone for p>.10.

Change N, or population effect size, and see generally lower or higher p values but, most surprisingly, in every case the values of p still jump around dramatically.

For videos of such dances, search YouTube for ‘dance of the p values’ and ‘significance roulette’.

Figures and dances like those shown here will come in Chapters 4, 5, and 6 in ITNS2.

Meanwhile, please have a play with Gordon’s wonderful dances and let us have your thoughts and suggestions. Thanks.

Geoff

The Shape of a Confidence Interval: Cat’s Eye or Plausibility Picture, and What About Cliff?

In brief:

  • Curves picture how likelihood varies across and beyond a CI. Which is better: One curve (plausibility picture) or two (cat’s eye)? Which should we use in ITNS2?
  • Curves can discourage dichotomous decision making based on a belief that there’s a cliff in strength of evidence at each limit of a 95% CI, i.e. at p=.05,
  • Explanation and familiarity are probably needed for curves on a CI to encourage estimation, rather than mere dichotomous interpretation.

Variation Across and Beyond a CI

A CI is most likely to land so that some point near the centre is at the unknown but fixed μ we wish to estimate. Less likely is that a point towards a limit is at μ. Of course there’s a 5% chance that a 95% CI lands so that μ is outside the interval. This pattern of relative likelihood is illustrated in Figure 1.2 from ITNS:

The curve illustrates the relative plausibility that various values along the axis are μ. The higher the curve, the better the bet that μ lies here. Keep in mind that our interval is one from the dance and that it’s the interval that varies over replication, while μ is assumed fixed but unknown. This single curve on a CI is the plausibility picture.

Cat’s Eyes

In this paper back in 2007 I played around with the black and white images at left, among others, as ways to picture how plausibility varies over and beyond the interval. The black bulge became the cat’s eye picture of a CI, as illustrated by the blue images in Figure 5.9 (below) from ITNS.

The 95% interval, in Figure 1.2 (above, at the top) and the middle of the blue figures, extends to include 95% of the area under the curve, or between the two curves. Similarly for the 99% and 80% CIs.

I don’t say that every graph with CIs needs to picture the cat’s eye, but do suggest that students and researchers would benefit from familiarity with the idea of plausibility changing smoothly across and beyond a CI. See any CI and, in your mind’s eye, see the cat’s eye bulge.

To what extent do researchers and students appreciate that pattern of variation across a CI? Pav Kalinowski and Jerry Lai, who worked with me years ago, investigated this question. This blog post (The Beautiful Face of a Confidence Interval: The Cat’s Eye Picture) describes their findings, with a link to the published results. In short, people’s intuitions are mostly inaccurate and highly diverse, but a bit of training and familiarity with the cat’s eye is encouragingly effective in improving their intuitions about the variation in plausibility. These results prompted use of the cat’s eye in UTNS (pp. 95-102) and ITNS.

Plausibility Pictures

More recently, the single curve rather than the two mirror image curves, has found favour: the plausibility picture, rather than the cat’s eye. Below is an example that Bob made in R, which appears in our eNeuro paper (Bob’s blog post is here). The CIs are 90% CIs.

The plausibility picture is shown here only on the CI on the difference, not on the two CIs to the left. It may be better to restrict the shading under the curve to the extent of the CI, and perhaps the curve could be half the height, so as not to be so visually dominant. Perhaps. Below is a variation on the same idea, from this preprint that Bob recently tweeted about.

The lower figure plots the mean and 95% CI, with plausibility picture, for the differences between the three rightmost conditions and the WT-EGFP condition at left. The dabestR package was used to make the figure. (The authors generously describe such a figure as a ‘Cumming estimation plot‘. I’m happy if UTNS and ITNS have popularised the use of a difference axis to picture a difference with its CI, but I later discovered that the idea goes back a while. The earliest examples I know of are in this 1986 BMJ article by Martin Gardner and Doug Altman, which includes two figures showing a difference with its CI on a difference axis, without any curve on that CI.) Let’s know of any earlier examples.

One or Two Curves? Plausibility Picture or Cat’s Eye?

Yes, plausibility picture or cat’s eye? I don’t know of any empirical study of which is more readily understood, or more effective in carrying the message of variation over the extent of a CI and beyond. There’s probably not much in it, so it comes down to a matter of taste. I’m sentimentally attached to the cat’s eye, but admit that the single curve is more visually parsimonious. Simplest would be a single fine line depicting the curve, with no shading. It would need to extend beyond both ends of the CI, but perhaps not by much. Perhaps such a curve is as good as anything. It would be great to have some evidence relevant to these questions. Meanwhile, I’d love to hear your views.

A Cliff at p=.05? At the End of a CI?

If a 95%CI is used merely to note whether or not the interval includes the null hypothesised value, we’re throwing away much of the information it offers and descending to mere dichotomous decision making. Undermining such ideas was one of my main motivations for playing with curves on a CI. In fact, no sharp change occurs exactly at a limit of a CI, as curves should make clear. Dropping the little crossbars at the end of the CI graphic (UTNS included crossbars, ITNS does not) was another attempt to de-emphasise CI limits.

To what extent do people think that a result falling just below p=.05 rather than just above it makes a difference? What about just inside or just outside a CI? Back in 1963 Rosenthal and Gaito, in this article (image below), asked psychology faculty and graduate students about their degree of confidence in a result, on a scale from 0 to 5, for various different p values. They identified a relatively steep drop in confidence either side of .05, and described this result as a cliff.

Here are their averaged results, degree of confidence plotted against p value:

Yes, the steepest part of the curves looks to be from .05 to .10, and results for the graduate students (top 3 curves) show a kink at .05, but the cliff is hardly precipitous. If we were not so indoctrinated about .05, perhaps we’d see these curves as suggesting a relatively steady drop, rather than a sudden cliff. I wonder whether any tendency towards cliff has increased since 1963?

Jerry Lai conducted an online version of this study, with published researchers from psychology and medicine as respondents. One version of his task asked about p values, another about CI figures–which showed intervals overlapping zero to varying extents corresponding to the various p values. His results are summarised here, along with brief mention of other similar studies since 1963. He found a diversity of shapes of curves: Only a few showed a steep cliff, many showed a weak cliff, as in the R & G average results above, some a more-or-less linear decline, and others some other shape. Psychology and medical researchers gave a similar diversity of curves. Results for CIs showed, if anything, more evidence of cliff than did those for p values. Alas!

Jerry’s chosen title for his article was: Dichotomous Thinking: A Problem Beyond NHST. In other words, CIs can easily be used merely to carry out NHST. A 2010 article from my group titled Confidence intervals permit, but do not guarantee, better inference than statistical significance testing reported evidence that researchers in psychology, behavioural neuroscience, and medicine tended to make much better interpretations of results shown with CIs if they avoided thinking about the CIs in terms of NHST.

A remarkable Open Science story

Jerry wondered how the curves for R & G’s individual participants may have varied from the average curves in the figure above. He wrote a very polite letter to Prof Rosenthal. By return of post came a charming and encouraging note to Jerry, enclosing several photocopied sheets of handwritten notes, which neatly set out full details of the experiment and the data for individuals. Yes, there was quite a diversity of curve shapes. Some 50 years on, the original data were still available! Bob Rosenthal was putting to shame many subsequent researchers who could not maintain data beyond the life of a particular computer and/or were not willing to share it with other researchers.

What About a Violin Plot?

I was delighted to see this preprint:

Bob tweeted about it a couple of weeks ago. It reports the results of online statistical cognition surveys. A blog post of ours a year ago, here, invited participation.

The authors asked participants to rate their confidence that an effect was non-zero, given CI figures corresponding to p values ranging from .001 through .04, .05, and .06, and up to .8. The figures included the standard 95% CI graphic, with little crossbars at the ends, and a violin plot as at left. The researchers found that they needed to give some explanation of the violin plot, especially considering that a violin plot usually represents the spread of data points, rather than a CI: it’s usually a descriptive rather than an inferential picture, as here. I suspect that would have been clearer if the violin plot had included the standard CI graphic–as the cat’s eye and plausibility picture do.

Overall, there was a small-to-moderate cliff effect between the .04 and .06 figures. The cliff was rather smaller for the violin plot than for the standard CI graphic.

Conclusions

  • The studies mentioned above don’t give us strong or definitive conclusions; we need replications.
  • Curves probably help CI interpretation, especially by discouraging mere dichotomous decision making.
  • The plausibility picture, perhaps without shading, may be the simplest and most parsimonious choice.
  • Some training and familiarity with any picture that includes one or more curves may be needed for full effectiveness.
  • There’s lots of scope for valuable empirical studies, perhaps especially of the plausibility picture.

A final question: ‘plausibility picture‘, ‘plausibility curve‘, ‘likelihood curve‘, ‘relative likelihood curve‘, or what? What’s your preference and why?

I’d love to have comments on these issues, and, especially, suggestions for our CI strategies in ITNS2, the second edition we’re currently working on.

Geoff

The Multiverse! Dances, and More, From Pierre in Paris

Our Open Science superego tells us that we must preregister our data analysis plan, follow that plan exactly, then emphasise just those results as most believable. Death to cherry-picking! Yay!

The Multiverse

But one of the advantages of open data is that other folks can apply different analyses to our data, perhaps uncovering interesting things. What if we’d like to explore systematically a whole space of analysis possibilities ourselves, to give a fully rounded picture of what our research might be revealing?

The figure below shows (a) traditional cherry-picking–boo!, (b) OCD following of fine Open Science practice–hooray!, and (c) off-the-wall anything goes–hmmm.

That fig is from a recent (in fact, forthcoming) paper by Pierre in Paris and colleagues. The paper is here, and the reference is below at the end (Dragicevic et al., 2019). The abstract below outlines the story.

The Multiverse, Live

Pierre and colleagues not only discuss the multiverse idea in that paper, but here they give neat interactive tools that allow any reader of several example papers to do the exploration themselves. Hover the mouse, or click, to explore the outcome of different analyses.

EMARs

I suggest Sections 4 and 5 in the paper are especially worth reading. Section 4 discusses what’s called explorable multiverse analysis reports (EMARs), with a focus on mapping out just what a rich range of possibilities there often are for alternative analyses.

Then Section 5 grapples with the (large) practical difficulties of building, reviewing, and using EMARs, with the aim of increasing insight into research results. Cherry-picking risks need always to be at the forefront of our thinking. Preregistration of certain proposed uses of an EMAR could be possible, with possibly somewhat reduced cherry-picking risks.

Play Multiverse on Twitter

Matthew Kay, one of the team, gave a great overview in 8 posts to Twitter. See the posts, and a bunch of GIFs in action here.

Enjoy!

Geoff

Pierre Dragicevic, Yvonne Jansen, Abhraneel Sarma, Matthew Kay, Fanny Chevalier. Increasing the Transparency of Research Papers with Explorable Multiverse Analyses. CHI 2019 – The ACM CHI Conference on Human Factors in Computing Systems, May 2019, Glasgow, United Kingdom. 2019, <10.1145/3290605.3300295>.