Of course I had to wear my APS uniform. You can just see, over Simine’s shoulder, the stunning artwork, portions of which appear on the cover of all my books.
It was great to see Simine last night, and to hear about the latest from APS and the world of journal editing. Her great love is still journal editing–very fortunately for psychology and indeed all of science. But when she wants quiet time for writing she takes to the road, finding quiet corners to use her laptop in coffee shops or, I was surprised to hear, breweries!
This last week she has slowly driven down from Sydney, coffee shop by coffee shop, to spend time with Fiona Fidler and with the MetaMelb group at The University of Melbourne.
Plastic plate, after 30 years of trips through the dishwasher
She dropped in for a tuna pasta dinner, which grandkids Lucy and Zoe always request. Simine ate off this StatPlay classic picture of the green heap, long since renamed the mean heap.
These days you can explore the dances and many other goodies in esci web.
P.S. On a personal note, we recently had a busy month to mark my 80thbirthday (yikes!)–wonderful chamber music, two big parties, and now, to cap it off, a visit from Simine!
…as evidenced by this article from Brazil, which I’m delighted to see:
The article’s header
I salute Karen Grimmer, JECP co-editor, for publishing it, and for managing to make it Open Access. Karen happens to be a long-standing friend of mine who now continues the good work in ‘retirement’. I declare an interest: I was a referee for the ‘Beyond the p Value…’ article.
Note the innovative review, evaluation, and synthesis techniques developed by the authors. Here’s the Abstract:
Rationale
The p value has long been used as the primary criterion for statistical significance; however, its dichotomous interpretation has been increasingly criticized for oversimplifying uncertainty and distorting scientific inference, particularly in health and sports sciences.
Aims and Objectives
This study aimed to critically analyze the limitations of using the p value as the central criterion of statistical significance and to discuss more robust methodological alternatives for statistical inference.
Methods
A critical review was conducted using the PubMed/MEDLINE database covering the period from 2015 to 2025, complemented by citation tracking. Reviews, editorials, guidelines, and methodological essays that directly addressed the interpretation of p values and complementary metrics were included. A total of 46 articles were selected and evaluated using a self-developed critical appraisal checklist.
Results
Among the included studies, 38 (82.6%) explicitly criticized the isolated or dichotomous use of the p value, whereas eight adopted a more moderate position, supporting its use only when combined with confidence intervals, effect sizes, or Bayesian approaches. No article defended the p value as a standalone criterion for scientific decision-making. The most frequent recommendations involved abandoning the term “statistically significant,” prioritizing the estimation of effect magnitude and precision, and promoting the use of compatibility intervals, effect sizes, and Bayesian methods.
Conclusion
Overcoming the binary logic of p < 0.05 is essential to enhance transparency, reduce bias, and better align statistical practice with the scientific and clinical relevance of research findings, particularly in the health and sports sciences.
A Practical Summary
A particularly useful feature is a dot point summary of practical recommendations near the end:
Pose quantitative research questions (“to what extent…?”).
Report effect sizes with compatibility intervals as primary results.
Jiangang Xia is an enterprising professor at the University of Nebraska who, among many other things, teaches into China. He alerted me some years ago to the difficulty his students in China had accessing my videos because YouTube was blocked for them. So I mounted the videos at this OSF site.
As another enterprising step, Jiangang has recently been making a number of posts to LinkedIn. Below is one of these:
Associate Professor of Educational Administration at University of Nebraska–Lincoln
January 13, 2026
Geoff Cumming and the New Statistics: Estimation as a Way of Thinking
Some years only reveal their importance in hindsight. For me—and for statistics reform—2014 was one of those years.
That was the year I had just begun my academic career at the University of Nebraska–Lincoln. It was also the year Rex Kline visited UNL and delivered his keynote, “Hello, Statistics Reform.” (for Kline’s talk, please see my Post 2 for more details). At the time, I had no idea how much that visit would eventually shape my thinking, teaching, and research.
What I did not know then—and only came to appreciate much later—is that 2014 was also the year Geoff Cumming delivered his workshop “The New Statistics: Estimation and Research Integrity” at the APS Annual Convention in San Francisco.
I was completely unaware of that workshop.
In fact, I would remain largely unaware of Cumming’s work for several more years—even though it quietly passed through my academic life more than once.
Missed encounters (seen only in hindsight)
While preparing this post, I did something simple: I searched Geoff Cumming’s name in my old emails. That’s when I realized how often his work had crossed my path without fully registering.
In 2016, I co-chaired a dissertation in which the student cited Cumming’s influential article “The New Statistics: Why and How” (2014). At the time, neither of us fully grasped what the article was really asking us to reconsider. Although reform language appeared, the dissertation still relied on phrases like “marginally significant” for p values above .05—an indication that NHST logic remained firmly in place. Estimation had been encountered, but not yet learned as a way of thinking.
In 2017, I received an email from Routledge inviting faculty to request a desk copy of Introduction to the New Statistics. I didn’t request it. Another missed opportunity—one I only recognize now, looking backward.
These weren’t personal oversights so much as reflections of how deeply NHST was normalized in our training. Reform ideas were present, but the infrastructure for learning them—how to teach them and how to use them—was still thin.
From critique to alternative
It wasn’t until 2019, through Rex Kline’s work, that the larger picture finally came into focus for me. From there, I traced the reform movement backward—and Geoff Cumming’s role became unmistakable.
Cumming and his colleagues were not simply extending Jacob Cohen’s critique of null hypothesis significance testing. Cohen had already shown, powerfully, why NHST was flawed and had pointed toward alternatives. He also played a central role in the APA Task Force on Statistical Inference, whose report was released in 1999.
Unfortunately, Cohen passed away in 1998, before that work could be fully carried forward. When the APA’s 5th edition was published in 2001, many of the reform ideas were only partially adopted, and everyday research norms remained largely unchanged.
What Cumming did next was different.
Through decades of scholarship—culminating in The New Statistics—he articulated a coherent alternative centered on estimation rather than binary decisions, emphasizing effect sizes, confidence intervals, precision, uncertainty, and cumulative evidence. By the New Statistics, Cumming meant an estimation-centered approach to inference—focusing on effect sizes, confidence intervals, and uncertainty, and on combining evidence across studies through meta-analysis—rather than making binary decisions based on p values alone.
From alternative to institutionalization and teaching
That work did not remain theoretical.
In 2008, Cumming was invited to join the small working group responsible for statistical reporting standards in the APA’s 6th edition, where he was a driving force behind the requirement that effect sizes and confidence intervals be reported for every research question—helping move estimation from an optional supplement to a core reporting standard.
Just as importantly, Cumming took the initiative to teach this alternative—by writing new articles (2014), new textbooks (2013, 2017, 2024), offering workshops (2014 APS), and developing demonstrations (ESCI) aimed at broad audiences, not just methodologists.
Why experience matters in teaching reform
Along the way, Cumming recognized something more fundamental: logical arguments alone were not enough.
As he later reflected, perhaps NHST had become “the researcher’s heroin—an addiction impervious to reason.”
If that was true, persuasion would require more than explanation. It would require experience.
This insight shaped his teaching. Rather than debating p values in the abstract, Cumming showed researchers—often viscerally—how unstable they are through demonstrations such as the dance of the p values, p intervals, and significance roulette. The goal was not just to convince the mind, but to engage the gut—to help researchers feel uncertainty rather than deny it.
A decade later, the field itself began to catch up. In 2024, Cumming’s 2014 article “The New Statistics: Why and How” received SAGE’s 10-Year Impact Award, recognizing research whose influence endures well beyond the standard citation window. The award was a reminder that reform ideas are often understood slowly—resisted early, adopted unevenly, and acknowledged only after they have quietly reshaped teaching and research norms.
Full circle: from missed encounters to transformed practices
What changed for me after 2019 was not just what I read—but how I taught, mentored, and conducted research.
In 2022, I shared Cumming’s 2014 article with my doctoral advisee Amanda. Her dissertation became the first I supervised to fully abandon significance language, adopting estimation-based interpretation throughout. That same year, a manuscript my student Cailen and I submitted—published in 2023—explicitly drew on Cohen, Kline, Cumming, and the ASA (2016) statement. It was my first research article grounded fully in estimation thinking.
And in a quiet but meaningful full circle, when I needed materials for my 2024 summer teaching in China, Geoff Cumming himself shared his 2014 APS workshop videos with me—materials I have since used both internationally and in my quantitative methods courses at UNL.
Looking back now, the reform was happening all around me in 2014.
I just wasn’t ready to see it.
Perhaps that is how intellectual change often works—not as a single revelation, but as a series of missed encounters that eventually align. And perhaps genuine reform requires not only better arguments, but better ways of helping researchers experience uncertainty.
Now he is close to completing Statistical Thinking for the AI Era, which comprises Vol. 1: Foundations and Inference, and Vol. 2: Advanced Methods–each with about 15 chapters and 520 pages. The chapters I have seen indicate it’s a terrific textbook.
Miodrag has lived and worked on at least four continents and takes a wholeheartedly global focus in both those major works. The textbook starts with Forewords from many continents. Miodrag generously invited me to contribute a foreword from Australia–which appears below. As you see, I slipped in mention that I’m an Antarctic tragic, even if it’s a big claim that I speak also from that seventh continent!
Foreword From Australia and Antarctica
I’ve been an Antarctic tragic since I was a boy. I’ve participated in citizen science during two voyages south. I’ve observed glacial retreat and penguin colonies dying. I read about giant datasets analysed by AI-assisted statistical modelling of changes to sea ice, weather, and much else including terrifying tipping points that will determine catastrophic changes our children must endure.
I’ve given a well-received statistics talk to researchers at the Institute for Marine and Antarctic Studies in a building on the wharf where Antarctic supply ships berth. All around us were whale skulls, ancient sledges, and other memorabilia. I feel I can claim at least a little justification for offering this piece on behalf of the seventh continent, as well as speaking from Australia.
Statistics is about communication and therefore the province of psychologists. Researchers publish representations of their results and readers, ranging from researchers to politicians, policymakers, and ordinary people, draw conclusions from those representations. However, vast industries misrepresent research results as they seek to persuade us to support autocrats, eat unhealthy food, buy things we don’t want that trash the planet, and burn fossil fuels to destroy the chance our children will inherit a liveable world.
My students and I studied people’s misconceptions of basic statistical concepts. That was statistical cognition, which has now developed hugely into metascience—a wonderfully interdisciplinary field that investigates how research is done, and should be done to be more trustworthy. For example, it studies how paper mills use AI tools to generate fake manuscripts that desperate researchers buy then submit to journals. Quality journals use AI tools and much editorial expertise to try to weed out fakes before these criminally pollute the research literature—much of it on life-and-death issues.
Around 2014 Open Science emerged: replication is central; this requires meta-analysis to synthesise results; this requires point and interval estimates, and statistical significance is irrelevant and often damaging. My statistics textbooks advocate estimation and meta-analysis. The intro book, with Bob Calin-Jageman, integrates estimation and Open Science all through. The second edition (Routledge, 2024) includes Bob’s wonderful open source software that’s ideal for beginners and researchers aiming for good estimation and Open Science practices. At thenewstatistics.com is more information, also the significance roulette simulation, which dramatizes a highly misleading feature of p values that few researchers appreciate. The good news is that these new approaches are found accessible by students and are a delight to teach.
I conclude that everyone should have a basic understanding of evidence and statistics. This book is broad in scope, well-informed, and future-focussed. For example, it explains traditional, Bayesian, and randomisation frameworks for estimation, any of which can support Open Science. As I read, I’m often applauding the approaches Professor Lovrić has chosen. With this book he is making an enormous contribution to statistics education around the world.
Simon Gandevia, a giant in the research integrity community, assembled an international A-list of incredible speakers–plus me–for this Conference (IRIC) recently in Sydney.
Diversity of disciplines, but a common focus
There were more than 100 registrants, so the presentation room was usually packed. A striking feature was our enormous diversity of disciplines: medicine, law, neuroscience, statistics, psychology, journalism, education… and more. Our common focus was the integrity of research, Open Science, and meta-research (research on how research is done). Most presentations to the conference were either an alarming assessment of how the research literature has been polluted, or a (sometimes) encouraging evaluation of efforts to improve integrity.
We felt strongly about so many things, with emotions swinging wildly. Mainly we felt enormous pride in group cohesion across such a diversity of backgrounds, and conviction that our mission–the very trustworthiness of research–is enormously important.
Paper mills, fakery, and more
But then we’d despair when a speaker documented the extend of deadly pollution spreading in the scientific literature, from old-fashioned fakery to plagiarism and fake manuscripts sold by paper mills. Anyone wanting to bulk up their cv can buy an AI-generated manuscript and submit it to a dozen journals. The best journals will identify the fakery, these days using AI to help, but some ‘vanity’ journal will publish the manuscript after a bogus peer review process–of course for a hefty payment. What criminality! What waste, especially of the time and effort of expert editors!
The memorable conference dinner was held on the roof deck of the stunning new Museum of Contemporary Art. Champagne in hand, sun setting on the harbour bridge to the left, opera house to the right: what’s not to like?
Significance Roulette released to the world
I’m delighted to say my brief presentation was very well received. Most importantly, I released to the world for the first time Bradley Dean‘s Significance Roulette that you can download, save, and run in your browser. Or simply go to esci web and select the significance roulette menu item to see and run the roulette wheel.
From here you can download my powerpoint slides, also a pdf summary, and the .html file you can save to your computer and run in your browser to enjoy significance roulette.
Read more about the rationale for significance roulette here.
Excel 2003, the best version ever, was enshittified by MicroSoft in the 2007 version, which was way slower and dropped many wonderful animation facilities :-(. Even vast efforts would not get my great Significance Roulette simulation running in the new version.
Significance Roulette, after an initial study gives p = .01
It has taken me close to 20 years to build the following argument that leads to significance roulette, and which is summarised in the first half of our open access article Calin-Jageman & Cumming, 2024
For approaching a century numerous distinguished scholars, including philosophers of science, statisticians, and psychologists, have published cogent critiques of p values, significance testing, and how researchers across science use these.
Even so, a large proportion of researchers, teachers, journals, and granting bodies use null hypothesis significance testing (NHST)–often based on p < .05 or p < .01–as the standard for concluding whether or not an effect exists, whether or not a result is large or important. Despite this logic being wrong-headed in so many ways!
Devotion to NHST and p <.05 resembles an addiction–the researcher’s heroin. Rational argument is not sufficient to shake the addiction. Could a dramatic demonstration, perhaps persuading via the gut rather than the brain, shake this addiction?
A striking but little-known feature of p values is that they are highly unreliable–their sampling variability is astonishingly large. Replicate a study, exactly the same but with a new random sample, and expect to obtain replication pthat can take just about any value between 0 and 1!
Jerry Lai in his lovely PhD studies took three converging approaches to find that a large proportion of published researchers in psychology, medicine, and statistics severely under-estimate the amount of variability in the p value with replication.
My first demonstration of p value variability was the dance of the p values, the first video of which dates from 2009. Search YouTube for ‘dance of the p values’ to find several videos. You can also play with the dances in esci web.
My second approach arose from study of the probability distribution of replication p, the p value obtained in a replication. My highly-cited 2008 article has details.
I define the p interval as the 80% prediction interval for replication p. It’s astonishingly long! For example, if an initial study obtains p = .05, the p interval is (.0002, .65), meaning an 80% chance of p within that interval and fully a 10% chance it falls below .0002 and 10% above .65. After p = .01 the interval is (.00001, .41). After initial p = .001 (*** highly statistically significant) there is fully a 1 in 6 chance a replication does not even achieve p < .05! After p = .20 ns there is a 1 in 3 chance a replication finds p < .05!
I take the probability distribution of replication p following initial p = .01 and divide the area under the curve, which represents probability, into 38 equal areas. I label each area with the p value in the centre of the interval. I have 38 p values, a few large and many small and very small, which accurately represent that probability distribution.
I scatter those 38 p values randomly around the 38 bins of a roulette wheel. Simply click to spin, wait for a few moments, and see the replication p you might have obtained from a replication. Much faster and cheaper than the hassle of raising a grant, recruiting participants, hassling with ethics approval… and collecting data!
At Significance Roulette in esci web you can click between initial p of .05 and .01. With initial p = .01, replication p values tend a little smaller, of course. But the striking thing is how widely spread the p values are! The distributions are pictured to left of the wheel. For example, switch from .05 to .01 and note a slightly smaller number of deep blue, deep trombone sound, despairing figures for p > .1 ns and slightly more bright red, triumphant trumpet blast, elated figures for p < .001 ***. Click SPIN, and note your quickened heart beat, sweaty palms, and that you are holding your breath–will you be despairing or elated–and in only a few seconds you’ll know!
Will this approach to tackling addiction via the gut be more effective than the decades of argument addressed at the cortex?
Definitely worth joining on 18 December, even if for me it’s at 6.30am. Note the first three speakers also kindly gave generous endorsements of ITNS2, at the start of the book. OS leaders, for sure.
In 2025, SIPS will hold its 10th annual meeting! In celebration of this milestone, please join us on December 18, 2024 (see starting time in your time zone) for the warm-up event SIPS 2025 Pre-Conference Discussion: What went wrong? How can we do it better?, during which a few early reformers will discuss challenges and improvements in psychological science. This 90-minute virtual event is free and open to all. Our invited speakers are:
Simine Vazire, Professor at the University of Melbourne, co-founder of SIPS
Brian Nosek, Executive Director of the Center for Open Science, co-founder of SIPS
Dorothy Bishop, University of Oxford
Joseph Simmons, Professor at the University of Pennsylvania, co-founder of Data Colada
Leif Nelson, Professor at UC Berkeley, co-founder of Data Colada
Moderator: Balazs Aczel, ELTE, host of SIPS 2025 in Budapest, Hungary.
I strongly recommend this lovely paper – it is full of fascinating examples and references; it is the type of paper you’ll want to assign to all your incoming trainees.
This week I (Bob) was part of a workshop on the future of neuroscience education. I got to be part of a rockstar panel with Tari Tan (Harvard), Monica Linden (Brown), and Rosalind Segal (Harvard). The event was organized by Lique Coolen (Kent State).
Tari spoke about curricular goals for neuroscience, reviewing an SFN initiative to define core competencies for trainees at the undergraduate and graduate level. I didn’t know about this! Tari gave a great overview; it’s well worth checking out: https://www.sfn.org/careers/higher-education-and-training/core-competencies
Monica discussed the role of generative AI in the future of neuroscience education and gave lots of thoughtful examples and resources.
Rosalind discussed developing an internship program for the neuroscience *PhD* program at Harvard! Very thougtful and interesting way for grad students to explore the diverse career paths that can come out of doctoral training — but also raised lots of interesting issues (a stated goal of internships that don’t impede research seemed hard to balance; students seemed to express worries about losing faculty support if they expressed ‘wrong’ career goals… lots to think about!).
My talk was about the ideal statistics curriculum. My short answer was “There isn’t one!” — neuro is too diverse and there is not only one ‘right’ way to do statistical inference. Still, the field of neuroscience shows some evident difficulties making valid claims from data (on that note, check out Chen et al., 2024), and so at least thinking about some broad ideals for training seems like a useful exercise. I came up with these guiding principles:
Estimation thinking at the forefront
Teach using simulations — for exploration at first, but eventually for planning experiments before they are conducted
Quickly graduate from toy examples to real, complex data sets and projects that require critical thinking about Multiplicity and Interdependence, and teach robustness checking to help students validate their approaches.
Integrate Open Science throughout
Ensure training is not just about statistics, but about the ‘neglected factors’, including good design and measurement — strength of evidence and quality of research are about so much more than just the statistics generated, and improvements are about so much more than sample size.
Chen, D. (2022). STATISTICAL PRACTICE IN PRECLINICAL NEUROSCIENCES: IMPLICATIONS FOR SUCCESSFUL TRANSLATION OF RESEARCH EVIDENCE FROM ANIMALS TO HUMANS Committee Member.
2014 was exciting: In January, Psychological Science‘s new requirements introduced many psychologists to Open Science. In May the APS Convention in San Francisco was an exhilarating (for me, anyway) whirl of symposia, talks, and discussions about how Open Science should be shaped and advanced. APS made 6 videos of my workshop ‘The New Statistics: Estimation and Research Integrity’. (Or use tiny.cc/apstnsvideos)
“Don’t we LOVE our p values!” I tossed the large p away, then jumped on it. Alas, my jumping didn’t make it into any of the six videos.
I was recently told the link to my workshop .ppt slide files was dead. I updated it, so tiny.cc/geoffdocs now points to a new OSF page with those files.
That led me to revisit the videos. They seem to me to hold up well–maybe even still worth considering. My one suggestion is to change every ‘research integrity‘ into ‘Open Science‘. OS is a way better term and what I would have used even a few months later.
It was the big Psychological Science changes and my article The New Statistics: Why and How that led to the workshop and videos. (For the article: tiny.cc/tnswhyhow)
With the recent Sage award for that article, it’s maybe timely to revisit the videos and again make the .ppt files accessible.