…as evidenced by this article from Brazil, which I’m delighted to see:
The article’s header
I salute Karen Grimmer, JECP co-editor, for publishing it, and for managing to make it Open Access. Karen happens to be a long-standing friend of mine who now continues the good work in ‘retirement’. I declare an interest: I was a referee for the ‘Beyond the p Value…’ article.
Note the innovative review, evaluation, and synthesis techniques developed by the authors. Here’s the Abstract:
Rationale
The p value has long been used as the primary criterion for statistical significance; however, its dichotomous interpretation has been increasingly criticized for oversimplifying uncertainty and distorting scientific inference, particularly in health and sports sciences.
Aims and Objectives
This study aimed to critically analyze the limitations of using the p value as the central criterion of statistical significance and to discuss more robust methodological alternatives for statistical inference.
Methods
A critical review was conducted using the PubMed/MEDLINE database covering the period from 2015 to 2025, complemented by citation tracking. Reviews, editorials, guidelines, and methodological essays that directly addressed the interpretation of p values and complementary metrics were included. A total of 46 articles were selected and evaluated using a self-developed critical appraisal checklist.
Results
Among the included studies, 38 (82.6%) explicitly criticized the isolated or dichotomous use of the p value, whereas eight adopted a more moderate position, supporting its use only when combined with confidence intervals, effect sizes, or Bayesian approaches. No article defended the p value as a standalone criterion for scientific decision-making. The most frequent recommendations involved abandoning the term “statistically significant,” prioritizing the estimation of effect magnitude and precision, and promoting the use of compatibility intervals, effect sizes, and Bayesian methods.
Conclusion
Overcoming the binary logic of p < 0.05 is essential to enhance transparency, reduce bias, and better align statistical practice with the scientific and clinical relevance of research findings, particularly in the health and sports sciences.
A Practical Summary
A particularly useful feature is a dot point summary of practical recommendations near the end:
Pose quantitative research questions (“to what extent…?”).
Report effect sizes with compatibility intervals as primary results.
He won the Nobel Prize for Economics in 2002 for foundational work on behavioural economics that was joint with Amos Tversky, who died in 1996.
I have two particular reasons for thinking of him:
‘Law’ of Small Numbers
A misconception rather than a law, this was described in the famous article Kahneman and Tversky (1971): Even quantitatively literate psychology researchers were likely to grossly over-estimate the probability that a replication of a study that obtained p = .05 would itself be statistically significant. This was a very early example of statistical cognition, the field that has been my primary research interest these last 25 years or so.
Moreover, that demonstration of drastic under-estimation of the sampling variability of the p value helped prompt my development of the dance of the p values, p intervals (Cumming, 2008), and statistical roulette–all of which are attempts to dramatise the sampling variability of p. I have hoped that such dramatisations would help undermine researchers’ seeming addiction to p values that has persisted despite numerous cogent critiques over more than half a century.
Anne Treisman
Anne, a distinguished cognitive psychologist, during 1968-1971 supervised my DPhil research at Oxford. In 1978 she married Daniel Kahneman. They moved to North America and were together at Princeton for many years before her death in 2018.
She appears at left receiving from Obama the U.S. National Medal of Science in 2013.
I salute the memory of these two fine scientists from whom I’ve learned an enormous amount.
Geoff
Cumming, G. (2008). Replication and p Intervals: p values predict the future only vaguely, but confidence intervals do much better. Perspectives on Psychological Science, 3(4), 286–300. https://doi.org/10.1111/j.1745-6924.2008.00079.x
Tversky, A., & Kahneman, D. (1971). Belief in the law of small numbers. Psychological Bulletin, 76(2), 105–110. https://doi.org/10.1037/h0031322
“It will keep you awake at night!” said Fiona Fidler about the Correcting the Record session. I’d zoomed in to some of AIMOS 2022 (the Meta-science conference in Melbourne last week; my post is here) but had missed that session. The video (here) is now online. It’s very much worth watching. I don’t expect to sleep for days 🙁
Four alarming and impressive speakers are listed above and mentioned below. Or search the AIMOS program(here) for the speaker names for more.
Editor-in-Chief of Anaesthesia and Intensive Care. Several submitted manuscripts every week show clear signs of fraud. Some are readily identified with a careful read. Plagiarism is rife and can be hard to detect. Chasing fraud now takes a large part of his time. His suggested tools:
“About 30% of the RCTs in Women’s Health are fabricated” (Weep!) Fabricating authors can be hard to hold to account, and they can attack back. We need better ways to identify and counter fraud. Are the data true? It can be hard or impossible to obtain original data.
Fraud is common, alas. Paper mills are companies that generate and sell fake manuscripts. Much but not all fraud is fairly easy to spot. She interviewed researchers to gather signs of fraud and is working towards a simple tool to help identify fraud. Others are using machine learning to develop such a tool.
Catching fraudulent and plagiarised images (Western blots, photos of tumors…) Exposing fraud can be a very long and frustrating process. Publishers can be obstructive, legalistic. Her suggested tools:
Finally
Yes, I was horrified at the extent of the problem–an enormous criminal enterprise polluting our research literature–often involving life-or-death issues. Is meta-science, and promotion of Open Science practices, the most urgent and consequential field for all of science at present? Maybe yes.
That video again. I hope you sleep well… eventually,
Online and in Melbourne, 28-30 Nov. A vast spread of disciplines. From around the world, numerous young folks–and some oldies–speaking truth to age. Yes, it’s AIMOS 2022. The fourth AIMOS conference! See my posts after the first conference (2019, here) and third (2021, here).
Registration
Register for AIMOS 2022 here. For online only it’s free. Or come along live and the social events alone more than justify the modest cost.
The Program
It’s here–scroll down a bit. It’s still being updated, more goodies to come.
…even to being a key financial Minister in Federal Parliament.
After the recent Australian election, Labor took power. (Hooray!) The new Assistant Minister for Treasury is Andrew Leigh. He’s one of the small team responsible for all things economic and financial, at the heart of Government.
Leigh is a Harvard-trained economist and former professor of economics. In Randomistas he argues we should use randomised trials much more often to guide public policy. Imagine, policy guided by high quality evidence! He’s well aware of the replication crisis and Open Science practices needed for trustworthy research.
In his AIMOS talk he discussed replication, His best one-liner: “If at first you DO succeed, try, try and try again.” Hey, meta-analysis!
Leigh is quoted in a recent interview as saying that he’s probably “the biggest stats nerd” the Australian Bureau of Statistics has had as its minister in its 116-year history. (Nerd can be good!)
I recently posted about MRI Together. It was a great global zoomfest, and now videos of the talks are online.
The Videos
The YouTube site with all the videos is here, but it’s easiest to scan the program and click on any talk title to go to the video.
It’s clear that many in the MRI community have been working on Open Science issues for a while. It was great to see lots of talks about software for statistical analysis of scans, and the challenges of making analyses reproducible–and of finding good ways to make code and the highly complex data sets open, and readily usable by others.
I’ll mention just a few videos below.
An Intro
Cassandra Gould van Praag gives an engaging intro talk here. She encountered a range of OS issues when her PhD examiners required modifications to her thesis. She bravely tells the story of the work and revisions she had to undertake. She must have taken the lessons to heart: Her current job is to develop and promote the open science infrastructure of Oxford Neuroscience.
Two ‘p Values Suck’ Talks
Valentin Amrhein gave a rousing talk titled ‘P Values and the Replicability of Results‘. The video is here. He has some well-chosen graphics and some striking graphs, for example of the wide range of likely p values in various situations. You won’t be surprised to hear that I heartily agree with almost everything he said.
My talk was titled The New Statistics for Reproducible Science The video is here. My three take-home messages were:
p values are highly unreliable, never trust them
Adopt Open Science practices, planned analyses
esci, Bob’s new software for estimation, meta-analysis, teaching, and more, is available now (in beta), and gaining thousands of installs each month. It includes great graphs with confidence intervals. It’s a pleasure to teach with it.
Brian Nosek Rounds It Off
Talk title: Publishing and Sharing Open Science. The video is here. Brian does a typically neat and persuasive job, of course with lots of evidence.
Lots more gems to discover: Just browse the program.
It’s fabulous to see yet one more research field, MRI and fMRI, jumping on board with Open Science. The Workshop runs 13-17 December, 2021, and the site is here.
Recall the dead salmon? (tiny.cc/deadsalmon, ITNS p. 485.) Back then, in 2009, a common way to analyse fMRI data relied on p values for each of many thousands of voxels. Apply this analysis to a dead salmon shown two different types of pictures and find part of its (totally dead) nervous system lit up on the analysis screen! It was a massive Type I error, of course, caused by inadequate correction of p values given by the gazillion voxel comparisons. Make a more appropriate correction and all we see is noise. The study deserved its 2012 IgNobel Prize.
Analysis of MRI data has come a long way since 2009, partly prompted by the dead salmon. More recent discussions, for example here and here, include consideration of the basic Open Science issues, including p-hacking, cherry-picking, unplanned analyses, and lack of replication.
The Timetable (Program)
It’s here. Near the top, select your timezone. There are sessions around the 24 hours, grouped to fit the waking hours in Atlantic, Pacific, Indian, and Caribbean zones. It looks to me like a wonderfully broad take on Open Science and Reproducibility. I recognised only a few of the presenters, including:
Valentin Amrhein: P-values and the replicability of results
Brian Nosek: Publishing and Sharing Open Science
I’m giving a brief (20 min, including Q+A) talk: The new statistics for reproducible science in a session titled Study design and interpretation. The reproducibility crisis, running 11.00 to 13.00 on 15 Dec, those times being UTC+11, which includes Sydney and Melbourne. I’ll post my slides in due course.
Replicate a study and you are highly likely to get a very different p value. Scarily different. The sampling variability of p values is so great that no p value deserves our trust.
Yes, I know, that’s been my mantra for more than a decade, but Steve Lindsay has just given me an excuse to beat the drum again. This is his great article:
It’s clear that Lindsay (2020) (preprint here) has been written by a practical, practising researcher. For example, there’s advice about references to help any new arrival in your lab get up to Open Science speed, and discussion about developing a laboratory manual to help everyone adopt good, systematic OS practices. His seven steps are an operationalisation of many such practices.
Possibly broader and more detailed accounts of what OS needs were given by Asendorpf et al. (2013), and Munafò et al. (2017). Dorothy Bishop (2020) gave a particularly compelling account, with a focus on overcoming the cognitive biases that make adopting OS practices a challenge.
Dance of the p Values
To illustrate his discussion of variability of ESs over replications, Steve Lindsay uses a figure (below) of the dance of the CIs and dance of the p values.
Screenshot from ESCI for UTNS. The blue and red distributions depict control and experimental populations, whose means differ by Cohen’s δ = 0.5 (or 10 raw points). The 20 open blue and 20 open red circles were randomly sampled from those populations. Each solid green circle represents the result of a simulated experiment: A raw effect size (experimental mean minus control mean) from random draws from the two populations, with the 95% CI around that ES estimate. The topmost (most recent) solid green circle represents the difference between the solid blue and red circles, the group means. In that random draw, the difference between conditions was not statistically significant (p = .229, as shown at left, corresponding to that CI considerably overlapping the vertical H0 line), a type II error. Of the 25 simulated experiments shown here, three came out in the wrong direction because of random sampling error; seven experiments detected the effect (p < .05, CI fully to the right of the H0 line) and every one of those overestimated the size of the effect.
CIs Make Uncertainty Salient
In the figure it’s easy, with a bit of practice, to focus on a single CI, note where it falls in relation to the vertical line (H0 of zero difference), and eyeball the p value. (ITNS Chap 6 has handy guidelines.) The two dances are telling us the same story.
But, and it’s a huge ‘but’, any single CI—which is all we know in real research life—gives us information about how wide, how frenetic the dance is. It makes the uncertainty salient. In stark contrast, any single p value tells us virtually nothing about the dancing p values. Its single value, perhaps reported to three decimal places, gives a seductive, but illusory, sense of certainty, even though it could have been very different. Yep, p is not to be trusted!
Software for the Dance of the p Values
I used ESCI in Excel 2003 (the best version ever, imho!) for the original dance of the p values video, also for this later video. It ran beautifully quickly. Unfortunately, Excel 2007 runs very much more slowly, for example in ESCI for UTNS, 2007 version, the replications plod down the screen.
Happily, we now have esci web, thanks to Gordon Moore. The first component is dances. Click ‘?’, top right in left panel, to turn on tips. Then have a play. For the dance of the p values, click red ‘9’, bottom sub-panel. Turn on sound and use the slider to adjust the speed, which can be too fast. Enjoy!
p Values Unreliable in Virtually Any Situation
Have I chosen sample sizes and population ES, δ, to give a particularly dramatic dance of the p values? You can use ESCI for UTNS, or esci web, to investigate. You will discover, for example, that larger sample sizes and/or larger δ, give average p that’s smaller (higher power, so more replications give p < .05), but there is still crazy dancing of the values of p.
For virtually any situation, the replication to replication variation in p is scarily large. Whenever you see a p value reported, remind yourself that the study could easily have given a quite different value, and that an exact replication is likewise likely to give a very different value. Any p value should be regarded as very fuzzy indeed.
What if We Don’t Know the Population ES, But Only the p Value?
The dance of the p values assumes that we know the population ES, δ, and N. However, in real research life we never know δ. In Cumming (2008) I developed the formulas for the sampling distribution of the p value for two situations:
δ and sample size(s) are known, as in the dances of the p values above; and
all we know is the p value given by an initial study.
For 2, we assume a single replication, everything exactly the same as in the initial study except that a new sample is taken. (Also assumes sample sizes are not small.) I refer to the p value given by such a replication as replication p.
The article includes figures showing the two distributions for a range of cases. For the second case, the figure shows the distribution of replication p following a specified p value found in the initial study.
To my surprise, I found that it’s perfectly possible to find and picture the distribution of replication p: Tell me only p from your study, and I’ll tell you the chance that an exact replication will give, for example, p < .05. Whatever the power, the N (if not small), and the true ES.
Not surprisingly, the distribution of replication p is very wide. Yes, following an initial p = .01 you are likely, on average, to get smaller replication p than following initial p = .05. But in both cases there’s vast uncertaintly—almost any replication p may occur.
Significance Roulette
How could I dramatize the almost unbelievably large amount of uncertainty in replication p? A mere diagram of the sampling distribution lacks punch. I divided the area under the distribution curve into 38 equal areas, because that’s the number of slots on one common roulette wheel. I used the p value at the centre of each area to represent that area. I arranged those 38 p values haphazardly around a wheel, in fact a roulette wheel. This figure is the wheel for initial p = .05:
After an initial study gives p = .05, a replication, just the same but with a new sample, will give replication p as shown by the wheel. The distribution of replication p is illustrated at left in terms of five conventional intervals. All p values are two-tailed.
If you find an initial p = .05, then to find what p an exact replication is likely to give, spin the wheel. Each of the 38 values is equally likely. At left is a summary, in terms of five conventional intervals of p values. You have a 7/38 chance of *** (p < .001, bright red circles) and 15/38 chance of p > .10 (deep blue circles). A mere tiny flip of the wheel and, instead of p = .39, as pictured, you might have obtained .02, or maybe ** or ***.
Here’s the wheel for initial p = .01:
As for the previous figure, but now for initial p = .01.
As we expect, replication p values are on average smaller than for the first wheel, but there’s still enormous variation. Chance of *** is now 12/38 and chance of p > .10 is now 10/38. Once again there’s vast uncertainty ☹
Videos for Significance Roulette
For better explanation than my brief sketch above, see two videos, here (tiny.cc/SigRoulette1) and here (tiny.cc/SigRoulette2). Incidentally, the wheel runs in Excel 2003 but, despite great effort, I haven’t been able to get it running in later Excel. If anyone would like to build it in some more modern language, preferably to run on the web, that would be great. Please let me know.
Beyond p Values
I’ve long argued, for example in Cumming (2014), that p values are rarely needed and that almost always we’re better off if we simply don’t report them, and don’t see them in published articles. They should simply be left to whither away. The only reason students should need to learn about them is to make sense of old research literature.
However, NHST and p values seem to be the researcher’s heroin. For most of us, and the published literature, they are deeply embedded and it seems very challenging to overcome the addiction. Rational argument has not proved effective. Can dramatization of unreliability do better?
“To What Extent…?”
For years I have been schooling myself never to ask “I wonder whether…” even when daydreaming. It always has to be “I wonder to what extent…”. Give it a try. That’s the step from dichotomous to estimation thinking. Our world simply isn’t black-white, but a gazzillion shades of grey, not to mention colours.
Meanwhile, enjoy seeing the wheel spin, and musing about what it tells us.
May all your confidence intervals be short!
Geoff
Asendorpf, J. B., et al. (2013). Recommendations for increasing replicability in psychology. European Journal of Personality, 27(2), 108-119. https://doi.org/10.1002/per.1919
Bishop, D. V. M. (2020). The psychology of experimental psychologists: Overcoming cognitive constraints to improve research: The 47th Sir Frederic Bartlett Lecture. Quarterly Journal of Experimental Psychology, 73(1) 1–19. https://doi.org/10.1177/1747021819886519
Cumming, G. (2008). Replication and p Intervals: p values predict the future only vaguely, but confidence intervals do much better. Perspectives on Psychological Science, 3(4), 286–300. https://doi.org/10.1111/j.1745-6924.2008.00079.x
Lindsay, D. S. (2020). Seven steps toward transparency and replicability in psychological science. Canadian Psychology/Psychologie canadienne, 61(4), 310–317. https://doi.org/10.1037/cap0000222
Concern about publication bias apparently goes back at least as far at Robert Boyle in the 17th century. That’s Barker Bausell’s starting point for his highly detailed account of the replication crisis and rise of Open Science. Bausell, the author of a number of books, brings a broad perspective, across many disciplines beyond only psychology, to the story he tells so engagingly. Frequent mentions of health and medical research reflect one aspect of his expertise.
Chapter 2 is appropriately critical in discussing NHST and p values, and false positives. Recommendations focus on using larger and higher-power studies, ESs, and α values much smaller than .05. I, of course, wish for recommendations against using p values at all, and in favour of estimation—for single studies and as the basis for meta-analysis. That’s surely at the core of Open Science practice.
Chapter 3 discusses a monumental list of 23 (!) QRPs—and some of those actually comprise more than one practice.
All through there’s much that will be familiar to many readers of this blog. We come across HARKing, forking paths, and Bem, Bargh, and Cuddy. But we also hear about cold fusion, and other goodies (?) from chemistry, physics, and more. I love the whole-of-science perspective. Of course, there’s lots of discussion of replication. Also of publishing practices, and the terrible contingencies that researchers have traditionally faced.
For any particular problem there’s discussion of the full sequence: the problem’s origins, investigations of how serious and widespread it is, what’s needed to solve it, then finally how we’re going so far with implementing the necessary changes. This sequence reminds me of Replicability, Robustness, and Reproducibility in Psychological Science, by Brian Nosek et al., which is a detailed discussion of exactly what the title states and developments over the last decade or so. This is in press for The Annual Review of Psychology and the final version is available here as a preprint. By comparison, Bausell ranges over many disciplines.
Education
Bausell’s brief final chapter looks forward. I like his emphasis on the incorporation of Open Science practices into education—I would say that, wouldn’t I, with ITNS2 designed to be a resource for just that. He takes an optimistic position, judging that, given sufficient support by researchers, positive developments in research practices are likely to flourish, even if success is not (yet) assured. Good to see; I dearly hope he’s correct.
While discussing the challenge of overcoming ingrained beliefs that ‘statistically significant’ is equivalent to ‘notable’, Bausell writes “the best educational option is probably to keep the sanctity of obtaining p-values at all costs from being learned in the first place” (p. 264). I couldn’t agree more! Again, I’d go further and advocate starting with estimation, then mentioning p values, if at all, only later.
What’s missing?
While Bausell generally gives us an upbeat and forward-looking discussion, I confess that I miss the sense of excitement that I’ve sometimes felt, for example at a couple of APS Conventions, roughly around 2014-2015. At those, some talks and symposia sparked animated discussions and a real sense that right here we are trying to figure out how science could be done better.
I also have in mind the youth of many advocates of better practices. I’ve had young researchers approach me after talks, sometimes in tears. They know what’s needed, but how, they ask, should they cope with a professor who says “reanalyse—your job is to find statistical significance in your data”? Despite such pressures, the Society for the Improvement of Psychological Science (SIPS) has always been energised by young folks. One-third of its board positions are reserved for graduate students and postdocs. I continue to be inspired and excited by young researchers who are insisting that science is done better, and who are active in helping figure out just how.
My conclusion is that Bausell gives us a valuable, wide-ranging and well-informed history, progress report, and agenda for the Open Science project.
The APS has just given a kick along to what’s most likely a myth: The Lucky Golf Ball. Alas!
Golf.com recently ran a story titled ‘Lucky’ golf items might actually work, according to study. The story told of Tiger Woods sinking a very long putt to send the U.S. Open to a playoff. “That day, Tiger had two lucky charms in-play: His Tiger headcover, and his legendary red shirt.”
The story cited Damisch et al. (2010), published in Psychological Science, as evidence the lucky charms may have contributed to the miraculous putt success.
Laudably, the APS highlights public mentions of research published in its journals. It posted this summary of the Golf.com story, and included it (‘Our science in the news’) in the latest weekly email to members.
However, this was a misfire, because the Damisch results have failed to replicate, and the pattern of results has prompted criticism of the work. Read on…
The original Lucky Golf Ball study
Damisch et al.reported a study in which students in the experimental group were told—with some ceremony—that they were using the lucky golf ball; those in the control group were not. Mean performance was 6.4 of 10 putts holed for the experimental group, and 4.8 for controls—a remarkable difference of d = 0.81 [0.05, 1.60]. (See ITNS, p. 171.) Two further studies using different luck manipulations gave similar results.
The replications
Bob and colleague Tracy Caldwell (Calin-Jageman & Caldwell, 2014) carried out two large preregistered close replications of the Damisch golf ball study. Lysann Damisch kindly assisted them make the replications as similar as possible to the original study. Both replications found effects close to zero.
Meta-analysis of original and replication studies
Here’s Figure 9.8 from ITNS, a forest plot that shows results from six studies from the Damisch group (red), and the two Calin-Jageman & Caldwell replications (blue). The first of the red studies and both of the blue used the lucky golf ball task.
The clear difference between the overall red mean and overall blue mean is shown at the bottom on a Difference axis; it’s -0.77 [-1.15, -0.38], so a clear failure to replicate.
Not replicable, but citable
That’s the title of a 2018 post by Bob lamenting the common pattern of a striking original finding that continues to make waves, even while strong counter evidence from replications languishes in the shadows.
Below I’ve updated his figure showing citation counts for the original Damisch article and the Bob & Tracy replication article. The pattern has not improved these last three years!
The pattern of Damisch results
The six red CIs in the forest plot are astonishingly consistent, all with p values a little below .05. Greg Francis, in this 2016 post, summarised several analyses of the patterns of results in the original Damisch article. All, including p-curve analysis, provided evidence that the reported results most likely had been p-hacked or selected in some way.
Another failure to replicate
Dickhäuser et al. (2020) reported two high-powered preregistered replications of a different one of the original Damisch studies, in which participants solved anagrams. Both found effects close to zero.
All in all, there’s little evidence for the lucky golf ball. APS should skip any mention of the effect.
What next?
Open Science practices will help. Perhaps high quality replication articles can be marked with big badges and trumpet fanfares? With everything online, it should be possible to add annotations to original articles and provide links to later replications, no doubt with original authors having a right of reply. Meanwhile, we need to keep up the skepticism and eternal vigilance.
Geoff
Calin-Jageman, R. J., & Caldwell, T. L. (2014). Replication of the Superstition and Performance Study by Damisch, Stoberock, and Mussweiler (2010). Social Psychology, 45(3), 239–245. https://doi/10.1027/1864-9335/a000190
Damisch, L., Stoberock, B., & Mussweiler, T. (2010). Keep Your Fingers Crossed! Psychological Science, 21(7), 1014–1020. https://doi/10.1177/0956797610372631
Dickhäuser, O., Heinze, A., Hamm, M. L., Bales, A. S., Bellmann, S. A., Böger, D., et al. (2020). Zwei teststarke, präregistrierte Replikationsstudien zum Einfluss von Glück auf kognitive Leistung (Two high-powered preregistered replication studies on effects of superstition on cognitive performance). Zeitschrift für Pädagogische Psychologie, 34, 51-60. https://doi.org/10.1024/1010-0652/a000263