Excel 2003, the best version ever, was enshittified by MicroSoft in the 2007 version, which was way slower and dropped many wonderful animation facilities :-(. Even vast efforts would not get my great Significance Roulette simulation running in the new version.
Two videos use my Excel 2003 version to explain: https://tiny.cc/SigRoulette1 and https://tiny.cc/SigRoulette2 or simply search at YouTube for ‘significance roulette’.
Now Bradley Dean has built Significance Roulette to run in your browser–as part of esci web.

It has taken me close to 20 years to build the following argument that leads to significance roulette, and which is summarised in the first half of our open access article Calin-Jageman & Cumming, 2024
- For approaching a century numerous distinguished scholars, including philosophers of science, statisticians, and psychologists, have published cogent critiques of p values, significance testing, and how researchers across science use these.
- Even so, a large proportion of researchers, teachers, journals, and granting bodies use null hypothesis significance testing (NHST)–often based on p < .05 or p < .01–as the standard for concluding whether or not an effect exists, whether or not a result is large or important. Despite this logic being wrong-headed in so many ways!
- Devotion to NHST and p <.05 resembles an addiction–the researcher’s heroin. Rational argument is not sufficient to shake the addiction. Could a dramatic demonstration, perhaps persuading via the gut rather than the brain, shake this addiction?
- A striking but little-known feature of p values is that they are highly unreliable–their sampling variability is astonishingly large. Replicate a study, exactly the same but with a new random sample, and expect to obtain replication p that can take just about any value between 0 and 1!
- Jerry Lai in his lovely PhD studies took three converging approaches to find that a large proportion of published researchers in psychology, medicine, and statistics severely under-estimate the amount of variability in the p value with replication.
- My first demonstration of p value variability was the dance of the p values, the first video of which dates from 2009. Search YouTube for ‘dance of the p values’ to find several videos. You can also play with the dances in esci web.
- My second approach arose from study of the probability distribution of replication p, the p value obtained in a replication. My highly-cited 2008 article has details.
- I define the p interval as the 80% prediction interval for replication p. It’s astonishingly long! For example, if an initial study obtains p = .05, the p interval is (.0002, .65), meaning an 80% chance of p within that interval and fully a 10% chance it falls below .0002 and 10% above .65. After p = .01 the interval is (.00001, .41). After initial p = .001 (*** highly statistically significant) there is fully a 1 in 6 chance a replication does not even achieve p < .05! After p = .20 ns there is a 1 in 3 chance a replication finds p < .05!
- I take the probability distribution of replication p following initial p = .01 and divide the area under the curve, which represents probability, into 38 equal areas. I label each area with the p value in the centre of the interval. I have 38 p values, a few large and many small and very small, which accurately represent that probability distribution.
- I scatter those 38 p values randomly around the 38 bins of a roulette wheel. Simply click to spin, wait for a few moments, and see the replication p you might have obtained from a replication. Much faster and cheaper than the hassle of raising a grant, recruiting participants, hassling with ethics approval… and collecting data!
- At Significance Roulette in esci web you can click between initial p of .05 and .01. With initial p = .01, replication p values tend a little smaller, of course. But the striking thing is how widely spread the p values are! The distributions are pictured to left of the wheel. For example, switch from .05 to .01 and note a slightly smaller number of deep blue, deep trombone sound, despairing figures for p > .1 ns and slightly more bright red, triumphant trumpet blast, elated figures for p < .001 ***. Click SPIN, and note your quickened heart beat, sweaty palms, and that you are holding your breath–will you be despairing or elated–and in only a few seconds you’ll know!
Will this approach to tackling addiction via the gut be more effective than the decades of argument addressed at the cortex?
Best of luck at the p Value Casino!
Take-home messages
- p values are unbelievably unreliable
- Any p value could easily have been just about any other value
- No p value deserves our trust 🙁
- Simply don’t use p values, there are much better ways 🙂
Geoff
P.S. Enormous thanks to Bradley and Bob, who made it all happen.













