Thankyou Jiangang! A Ten-Year Journey to Significance Roulette

Jiangang Xia is an enterprising professor at the University of Nebraska who, among many other things, teaches into China. He alerted me some years ago to the difficulty his students in China had accessing my videos because YouTube was blocked for them. So I mounted the videos at this OSF site.

As another enterprising step, Jiangang has recently been making a number of posts to LinkedIn. Below is one of these:

From a LinkedIn post by Jiangang Xia

Rethinking Quantitative Reasoning in Educational Research (4)

Jiangang Xia

Jiangang Xia

Associate Professor of Educational Administration at University of Nebraska–Lincoln

January 13, 2026

Geoff Cumming and the New Statistics: Estimation as a Way of Thinking

Some years only reveal their importance in hindsight. For me—and for statistics reform—2014 was one of those years.

That was the year I had just begun my academic career at the University of Nebraska–Lincoln. It was also the year Rex Kline visited UNL and delivered his keynote, “Hello, Statistics Reform.” (for Kline’s talk, please see my Post 2 for more details). At the time, I had no idea how much that visit would eventually shape my thinking, teaching, and research.

What I did not know then—and only came to appreciate much later—is that 2014 was also the year Geoff Cumming delivered his workshop “The New Statistics: Estimation and Research Integrity” at the APS Annual Convention in San Francisco.

I was completely unaware of that workshop.

In fact, I would remain largely unaware of Cumming’s work for several more years—even though it quietly passed through my academic life more than once.

Missed encounters (seen only in hindsight)

While preparing this post, I did something simple: I searched Geoff Cumming’s name in my old emails. That’s when I realized how often his work had crossed my path without fully registering.

In 2016, I co-chaired a dissertation in which the student cited Cumming’s influential article “The New Statistics: Why and How” (2014). At the time, neither of us fully grasped what the article was really asking us to reconsider. Although reform language appeared, the dissertation still relied on phrases like “marginally significant” for p values above .05—an indication that NHST logic remained firmly in place. Estimation had been encountered, but not yet learned as a way of thinking.

In 2017, I received an email from Routledge inviting faculty to request a desk copy of Introduction to the New Statistics. I didn’t request it. Another missed opportunity—one I only recognize now, looking backward.

These weren’t personal oversights so much as reflections of how deeply NHST was normalized in our training. Reform ideas were present, but the infrastructure for learning them—how to teach them and how to use them—was still thin.

From critique to alternative

It wasn’t until 2019, through Rex Kline’s work, that the larger picture finally came into focus for me. From there, I traced the reform movement backward—and Geoff Cumming’s role became unmistakable.

Cumming and his colleagues were not simply extending Jacob Cohen’s critique of null hypothesis significance testing. Cohen had already shown, powerfully, why NHST was flawed and had pointed toward alternatives. He also played a central role in the APA Task Force on Statistical Inference, whose report was released in 1999.

Unfortunately, Cohen passed away in 1998, before that work could be fully carried forward. When the APA’s 5th edition was published in 2001, many of the reform ideas were only partially adopted, and everyday research norms remained largely unchanged.

What Cumming did next was different.

Through decades of scholarship—culminating in The New Statistics—he articulated a coherent alternative centered on estimation rather than binary decisions, emphasizing effect sizes, confidence intervals, precision, uncertainty, and cumulative evidence. By the New Statistics, Cumming meant an estimation-centered approach to inference—focusing on effect sizes, confidence intervals, and uncertainty, and on combining evidence across studies through meta-analysis—rather than making binary decisions based on p values alone.

From alternative to institutionalization and teaching

That work did not remain theoretical.

In 2008, Cumming was invited to join the small working group responsible for statistical reporting standards in the APA’s 6th edition, where he was a driving force behind the requirement that effect sizes and confidence intervals be reported for every research question—helping move estimation from an optional supplement to a core reporting standard.

Just as importantly, Cumming took the initiative to teach this alternative—by writing new articles (2014), new textbooks (2013, 2017, 2024), offering workshops (2014 APS), and developing demonstrations (ESCI) aimed at broad audiences, not just methodologists.

Why experience matters in teaching reform

Along the way, Cumming recognized something more fundamental: logical arguments alone were not enough.

As he later reflected, perhaps NHST had become “the researcher’s heroin—an addiction impervious to reason.”

If that was true, persuasion would require more than explanation. It would require experience.

This insight shaped his teaching. Rather than debating p values in the abstract, Cumming showed researchers—often viscerally—how unstable they are through demonstrations such as the dance of the p values, p intervals, and significance roulette. The goal was not just to convince the mind, but to engage the gut—to help researchers feel uncertainty rather than deny it.

A decade later, the field itself began to catch up. In 2024, Cumming’s 2014 article “The New Statistics: Why and How” received SAGE’s 10-Year Impact Award, recognizing research whose influence endures well beyond the standard citation window. The award was a reminder that reform ideas are often understood slowly—resisted early, adopted unevenly, and acknowledged only after they have quietly reshaped teaching and research norms.

Full circle: from missed encounters to transformed practices

What changed for me after 2019 was not just what I read—but how I taught, mentored, and conducted research.

In 2022, I shared Cumming’s 2014 article with my doctoral advisee Amanda. Her dissertation became the first I supervised to fully abandon significance language, adopting estimation-based interpretation throughout. That same year, a manuscript my student Cailen and I submitted—published in 2023—explicitly drew on Cohen, Kline, Cumming, and the ASA (2016) statement. It was my first research article grounded fully in estimation thinking.

And in a quiet but meaningful full circle, when I needed materials for my 2024 summer teaching in China, Geoff Cumming himself shared his 2014 APS workshop videos with me—materials I have since used both internationally and in my quantitative methods courses at UNL.

Looking back now, the reform was happening all around me in 2014.

I just wasn’t ready to see it.

Perhaps that is how intellectual change often works—not as a single revelation, but as a series of missed encounters that eventually align. And perhaps genuine reform requires not only better arguments, but better ways of helping researchers experience uncertainty.

For readers who want to explore further

(All materials shared with permission; enormous credit to Geoff Cumming, Bradley Dean, and Robert Calin-Jageman.)

Dear colleagues,

  • When did you first encounter Geoff Cumming’s work—if at all?
  • Have you seen estimation treated as an add-on, rather than a way of thinking?
  • What important ideas did you meet early, but only understand much later?

#Cumming #NewStatistics #EstimationThinking #QuantitativeMethods #EducationalResearch

Thank you Jiangang!

Geoff

The p Value Casino Is Open–For Significance Roulette!

Excel 2003, the best version ever, was enshittified by MicroSoft in the 2007 version, which was way slower and dropped many wonderful animation facilities :-(. Even vast efforts would not get my great Significance Roulette simulation running in the new version.

Two videos use my Excel 2003 version to explain: https://tiny.cc/SigRoulette1 and https://tiny.cc/SigRoulette2 or simply search at YouTube for ‘significance roulette’.

Now Bradley Dean has built Significance Roulette to run in your browser–as part of esci web.

Significance Roulette, after an initial study gives p = .01

It has taken me close to 20 years to build the following argument that leads to significance roulette, and which is summarised in the first half of our open access article Calin-Jageman & Cumming, 2024

  • For approaching a century numerous distinguished scholars, including philosophers of science, statisticians, and psychologists, have published cogent critiques of p values, significance testing, and how researchers across science use these.
  • Even so, a large proportion of researchers, teachers, journals, and granting bodies use null hypothesis significance testing (NHST)–often based on p < .05 or p < .01–as the standard for concluding whether or not an effect exists, whether or not a result is large or important. Despite this logic being wrong-headed in so many ways!
  • Devotion to NHST and p <.05 resembles an addiction–the researcher’s heroin. Rational argument is not sufficient to shake the addiction. Could a dramatic demonstration, perhaps persuading via the gut rather than the brain, shake this addiction?
  • A striking but little-known feature of p values is that they are highly unreliable–their sampling variability is astonishingly large. Replicate a study, exactly the same but with a new random sample, and expect to obtain replication p that can take just about any value between 0 and 1!
  • Jerry Lai in his lovely PhD studies took three converging approaches to find that a large proportion of published researchers in psychology, medicine, and statistics severely under-estimate the amount of variability in the p value with replication.
  • My first demonstration of p value variability was the dance of the p values, the first video of which dates from 2009. Search YouTube for ‘dance of the p values’ to find several videos. You can also play with the dances in esci web.
  • My second approach arose from study of the probability distribution of replication p, the p value obtained in a replication. My highly-cited 2008 article has details.
  • I define the p interval as the 80% prediction interval for replication p. It’s astonishingly long! For example, if an initial study obtains p = .05, the p interval is (.0002, .65), meaning an 80% chance of p within that interval and fully a 10% chance it falls below .0002 and 10% above .65. After p = .01 the interval is (.00001, .41). After initial p = .001 (*** highly statistically significant) there is fully a 1 in 6 chance a replication does not even achieve p < .05! After p = .20 ns there is a 1 in 3 chance a replication finds p < .05!
  • I take the probability distribution of replication p following initial p = .01 and divide the area under the curve, which represents probability, into 38 equal areas. I label each area with the p value in the centre of the interval. I have 38 p values, a few large and many small and very small, which accurately represent that probability distribution.
  • I scatter those 38 p values randomly around the 38 bins of a roulette wheel. Simply click to spin, wait for a few moments, and see the replication p you might have obtained from a replication. Much faster and cheaper than the hassle of raising a grant, recruiting participants, hassling with ethics approval… and collecting data!
  • At Significance Roulette in esci web you can click between initial p of .05 and .01. With initial p = .01, replication p values tend a little smaller, of course. But the striking thing is how widely spread the p values are! The distributions are pictured to left of the wheel. For example, switch from .05 to .01 and note a slightly smaller number of deep blue, deep trombone sound, despairing figures for p > .1 ns and slightly more bright red, triumphant trumpet blast, elated figures for p < .001 ***. Click SPIN, and note your quickened heart beat, sweaty palms, and that you are holding your breath–will you be despairing or elated–and in only a few seconds you’ll know!

Will this approach to tackling addiction via the gut be more effective than the decades of argument addressed at the cortex?

Best of luck at the p Value Casino!

Take-home messages

  • p values are unbelievably unreliable
  • Any p value could easily have been just about any other value
  • No p value deserves our trust 🙁
  • Simply don’t use p values, there are much better ways 🙂

Geoff

P.S. Enormous thanks to Bradley and Bob, who made it all happen.

Open Science: Free Zoom With the Experts Next Week

Definitely worth joining on 18 December, even if for me it’s at 6.30am. Note the first three speakers also kindly gave generous endorsements of ITNS2, at the start of the book. OS leaders, for sure.

Info and registration here. The announcement:

In 2025, SIPS will hold its 10th annual meeting! In celebration of this milestone, please join us on December 18, 2024 (see starting time in your time zone) for the warm-up event SIPS 2025 Pre-Conference DiscussionWhat went wrong? How can we do it better?, during which a few early reformers will discuss challenges and improvements in psychological science. This 90-minute virtual event is free and open to all. Our invited speakers are:

  • Simine Vazire, Professor at the University of Melbourne, co-founder of SIPS 
  • Brian Nosek, Executive Director of the Center for Open Science, co-founder of SIPS 
  • Dorothy Bishop, University of Oxford
  • Joseph Simmons, Professor at the University of Pennsylvania, co-founder of Data Colada
  • Leif Nelson, Professor at UC Berkeley, co-founder of Data Colada

Moderator: Balazs Aczel, ELTE, host of SIPS 2025 in Budapest, Hungary.  

CLICK HERE TO REGISTER You can find out more about the event at https://improvingpsych.org/sips-2025-precon-discussion (including a link to submit your questions in advance).

ITNS2 Data Sets Are Now Available Within esci

All the data sets used in examples and exercises in the new second edition of ITNS are now available within esci in jamovi, esci in R, and soon within esci in JASP.

Here’s how to open one of those data sets in jamovi.

  1. Open jamovi with the esci module installed–meaning you can see the esci icon in the top bar, as pictured.
  2. Click three lines, top left. Panel opens.
  3. Click Open, Data Library.
  4. See 47 data files, ordered by chapter of first mention in the book.
  5. Click on your choice of data set.

In the book, ignore instructions to go to the book’s Companion Website to download a data set and save locally before opening in jamovi.

Thanks Bob, life is now easier!

Geoff

‘The New Statistics’ (2013) Wins Sage 10-Year Impact Award

The New Statistics: Why and How (abstract below) explained the advantages of moving on from NHST to the new statistics (estimation and meta-analysis) and the need for better practices to improve research integrity. I’m delighted that an award from Sage indicates the article seems to be helping researchers improve what they do. Next: Can ITNS2 help the next generation do even better?

The article appeared online in late 2013, so was considered when Sage examined the citation numbers of all articles appearing in any of the 400+ journals Sage published back in 2013. It was one of the top three most cited, so has been given a Sage 10-Year Impact Award. Sage’s announcement is here. Sage has just published a blog post about it with headline:

Now for the abstract:

The article was commissioned by Eric Eich, then editor-in-chief of Psychological Science, to appear immediately following his famous editorial Business Not As Usual. This opened the Journal’s first issue of 2014 and announced sweeping changes in the journal’s submission requirements, which, for many psychologists, marked the arrival of Open Science.

Interview

Sage’s blog post includes an email interview with me. Here’s a brief summary:

What was it in your own background that led to your article?

When I was a teenager my father gave me a simple explanation of significance testing. I said something like “That’s weird, sort of backwards. And why .05?” He replied “I agree, but that’s the way we do it.”

Over decades of teaching I became ever more dissatisfied with NHST, and focused ever more on confidence intervals (CIs).

Was there an article that had a particularly strong influence on you?

Frank Schmidt (1996) wrote: “It is now possible to use meta-analysis to show that reliance on significance testing retards the development of cumulative knowledge.” A revelation!

What did Schmidt’s article lead to?

About 2003 I started using an Excel forest plot to give a simple explanation of meta-analysis in my intro course. I was delighted: Students told me it just made sense. Of course, for meta-analysis you need a CI from each study, while p values are irrelevant, even misleading.

      Figure: Dances of means, confidence intervals, and p values.

In 2009 I uploaded a video of the dance of the p values. I became passionate about advocating the new statistics (estimation and meta-analysis). I wrote Understanding The New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis (UTNS, 2012).

What was happening in psychology at about that time?

Ioannidis (2005) explained how reliance on NHST was a major cause of the replication crisis. Largely in response to that crisis, Open Science arrived—perhaps the most important advance in how science is done for a very long time.

Eric Eich’s famous editorial Business Not As Usual in the January 2014 issue of Psychological Science marked the arrival of Open Science in psychology. Months earlier Eric had invited me to write a tutorial article to support the changes he wanted. This was The New Statistics: Why and How and was published immediately following his editorial.

What has been the reception of the article?

Mainly very positive. Some have felt I went too far in advising that in most cases it’s better not to use NHST at all. Some Bayesians have been unhappy with the focus on confidence intervals.

Revisiting that article, what would you have done differently?

I used the term ‘research integrity’, but ‘Open Science’ was coming into use and I soon realized that was way better. Reading the article today, for ‘research integrity’ read ‘Open Science’.

Otherwise, I think the article has held up well, including all 25 guidelines in Table 1.

What has happened since

Psychological Science has continued to lead in the adoption of Open Science practices.

Meta-science, also known as meta-research, has emerged and now thrives as a highly multi-disciplinary field. It applies the scientific method to improve that method—wonderful!

What have you been doing since?

I teamed with Robert Calin-Jageman to write the first intro statistics textbook based on the new statistics and with Open Science all through. The second edition has just come out: Introduction to The New Statistics: Estimation, Open Science, and Beyond, 2nd edition (ITNS2, 2024). It has much improved software, as we explain in Calin-Jageman & Cumming (2024), which is on open access.

We believe this book can sweep the world—we’ll see! To read the Preface and Chapter 1 go to www.thenewstatistics.com. In the second para is a link to the book’s Amazon site. Click ‘Read sample’.  

References

Calin-Jageman, R., & Geoff Cumming, G. (2024). From significance testing to estimation and Open Science: How esci can help. International Journal of Psychology,     https://doi.org/10.1002/ijop.13132

Cumming, G. (2012). The New Statistics: Effect sizes, confidence intervals, and meta-analysis. New York: Routledge. 

Cumming, G. (2014) The new statistics: Why and how. Psychological Science. 25(1), 7-29. https://doi.org/10.1177/0956797613504966

Cumming, G., & Calin-Jageman, R. (2024). Introduction to The New Statistics: Estimation, Open Science, & Beyond, 2nd edition. New York: Routledge.

Eich, E. (2014) Business not as usual. Psychological Science, 25(1), 3–6. https://doi.org/10.1177/0956797613512465

Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine 2: e124. https://doi.org/10.1371/journal.pmed.0020124

Schmidt, F. L. (1996). Statistical significance testing and cumulative knowledge in psychology: Implications for training of researchers. Psychological Methods, 1(2), 115-129. https://doi.org./10.1037/1082-989X.1.2.115 

Booklisti. For Finding Interesting Books, Now Including ITNS2

Exploring Booklisti is a neat way to find good books to read.

Booklisti, would you believe, comprises lots of short lists of books that hang together. I have two lists, ITNS2 appearing in each. My first list is just UTNS, my first book, and ITNS2, our second edition of the intro book.

My second list (go to that site for links to the books listed below), titled Open Science and how to do better research with better statistics, comprises:

Introduction to The New Statistics, Second Edition By Geoff Cumming and Robert Calin-Jageman

Understanding The New Statistics By Geoff Cumming

A Student’s Guide to Open Science By Charlotte Pennington

The Seven Deadly Sins of Psychology: A Manifesto for Reforming the Culture of Scientific Practice By Chris Chambers

Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth By Stuart Ritchie

Beyond Significance Testing By Rex B. Kline

Research Methods in Psychology: Evaluating a World of Information By Beth Morling

The Design of Experiments in Neuroscience By Mary E. Harrington

Happy exploring of Booklisti and happy reading.

Geoff

Meet Petra: Enthusiasm for Archaeology, Open Science, and Better Statistics

Lunch with Petra Vaiglova

It was a pleasure to meet Petra Vaiglova a few days ago while she was in Melbourne for an archaeological science conference. Fiona Fidler joined us for lunch–thanks to her for hosting.

Originally from the Czech Republic (Czechia), Petra has lived, studied, and worked all over the world, as you can see at her site. Her doctorate is from Oxford. Her TEDx talk outlines some of her research interests.

She arrived in Australia a couple of years ago as a post doc at Griffith University, Queensland. Soon after, she launched into organising what became a three-day online Workshop on Good Statistical Practice in Archaeology–open to anyone of any discipline.

I first learned of her enthusiasm for statistical reform and Open Science when she kindly invited me to speak at that workshop. I gave a talk on the new statistics, and another on using Bob’s new esci software in jamovi–in archaeological science, as in any other discipline. I posted about the workshop here.

Earlier this year Petra took up a lectureship at ANU in Canberra, and enthusiastically took on teaching statistics and topics in archaeological science to both undergraduate and postgraduate students.

She expressed keen interest in our second edition, even volunteering to help. Over many months last year she worked through final drafts of most chapters, picking up many errors and infelicities. Later she worked painstakingly through the proofs of many chapters, picking up elusive tiny errors. She made an immense contribution to ITNS2, as Bob and I acknowledge on p. xxvii.

So last week Petra, Fiona and I had much to discuss. Petra will be at AIMOS in November (more details here), no doubt meeting many of the good folks who make the Meta-science scene in Australia so lively and multi-disciplinary.

As an associate editor of the Journal of Archaeological Science she is helping develop guidelines for that journal to encourage reproducibility and Open Science practices.

I discovered years ago that archaeological science has Open Science lessons for us all–see my post here. I wish Petra all strength for her continuing efforts towards statistical reform and Open Science!

Geoff

To Find Interesting Books, Explore bookdna.com, Now Including ITNS2

A couple of years back I posted (here) about bookdna.com, which has gone from strength to strength as an engaging way to browse books for interesting finds. I’ve updated our bookdna entry to ITNS2 and tweaked our recommendations, with Pennington‘s little gem on Open Science now first on the list. The start of our entry:

Our recommended list is now:

A Student’s Guide to Open Science By Charlotte Pennington

Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth By Stuart Ritchie

Beyond Significance Testing By Rex B. Kline

Research Methods in Psychology: Evaluating a World of Information By Beth Morling

The Design of Experiments in Neuroscience By Mary E. Harrington

bookdna offers lots of ways to explore. Happy reading.

Geoff

Choosing a Textbook Cover Design

It’s a delicious moment when the publisher sends a number of options their graphic designer has dreamed up for the cover. Below are the options for the three books. In each case, can you pick our choice? Our choices are below—don’t scroll down yet… 

UTNS (2012) …at left.

ITNS1 (2017) …below.

ITNS2 (2024) …below.

The three sets, all framed expertly by Lindsay my wife, hang above my desk:

Our Choices …do you think we got it right?

UTNS (2012) …at left

Middle option

 ITNS1 (2017) …at right

Leftmost option

ITNS2 (2024) …below

Top right option, as below left. (We were offered just the other five but asked to see the bottom right design in the bottom left colours, so Routledge sent the top right option, which we chose.)

However, when we received our printed books, we discovered that Routledge had actually used a modification of our chosen design, as at right. Not exactly our choice, but not bad.

‘Treasure’: Claire’s Gorgeous Resin Artwork

Treasure, at left, by Claire Layman, 150 × 50cm, resin on stretched canvas. Claire is an internationally recognised artist, also a longtime friend.

Walk into our living room and be struck by the vibrancy and depth of colour of Treasure, so much more alive than any small printed copy can be.

Claire generously agreed that Treasure could be used on the cover of our three statistics textbooks. See the note on the copyright page of each.

At right, top to bottom:

UTNS (2012)

ITNS1 (2017)

ITNS2 (2024)

People often make comments like: “Looks like slices through some sort of stones”, “It’s wriggling things under a microscope!” or “Go snorkelling and see things like those?”

Claire’s response is “It’s an artwork, see it as you wish!” She mentions also that there’s no official top or bottom: hang it horizontally or vertically, either way up.

Besides looking great, is there any justification for it appearing on statistics books? People often make comments like: “There’s a pattern of those blues—oh, no there’s not”, or “Look, those ones sort-of alternate, but not quite”. Claire says that she often “starts to make a pattern, then breaks it”. That all sounds to me like trying to find some sort of regularity in randomness, which is one way of describing the aim of statistical inference: Can we identify a difference, or other pattern, lurking in the sampling variability, how large or strong is that pattern, and how confident can we be in our conclusion? That’s the central concern of our books.

Do you agree with the graphic designer’s choice of part-images from Treasure for the covers?

Geoff

NEXT: Choosing a cover design.