Brunch With Petra at the Museum

Petra Vaiglova is now Senior Lecturer in Archaeological Science at the Australian National University in Canberra. I posted here about meeting her for the first time, when she visited The University of Melbourne a few years back.

Geoff, Petra, and Stephen. Behind is a very early model Holden, manufactured in Australia, towing a caravan also of about 1950.

Since then she has published two especially notable open-access articles, the first being How can we improve statistical training in archaeological science?

Improving statistical training in archaeology science

Here’s the graphical abstract, the great work of Kathryn Killackey:

See the P.P.S. below for the full abstract.

Teeth, and ritual feasting a long long time ago

The second article, also open-access, is Transport of animals underpinned ritual feasting at the onset of the Neolithic in southwestern Asia, which reports a highly innovative study, led by Petra, of teeth from an Early Neolithic site in Iran.

I was recently in Canberra and, happily, could catch up with Petra and her partner, Stephen, for brunch at the National Museum cafe. I enjoyed a very good brunch, with animated and highly interesting discussion.

Geoff

P.S. Petra also very kindly made the trek from Canberra to join a rcent party marking my 80th birthday.

Yikes!

P.P.S. The abstract of Petra’s statistics article:

  • Raising the standard for statistical training in archaeology will improve the breadth and depth of archaeological science.
  • Improving statistical training can start by discussing five fundamental statistical concepts that archaeologists do not talk about enough.
  • Supervisors can help make statistical training more effective by advocating for statistical reform and Open Science.

The aim of this paper is to shine light on fundamental statistical concepts that archaeologists do not talk about enough. I argue that more deliberate discussion of these statistical ‘elephants in the room’ can have a positive impact on improving statistical training and on steering us away from perpetuation of poor research practices.

1) Statistical thinking should come first. This will help us break down some of the stigma around numbers and statistics, and set us up for building analytical frameworks that will provide the most informative answers to our research questions.

2) Descriptive and inferential statistics have different interpretative potential. This will clarify how we can move from using tools that only allow us to talk about our studied samples to using tools that enable us to draw inferences about the underlying populations from which the samples derived.

3) p values can be extremely variable. This will help spread awareness about the misuses and misconceptions of Null Hypothesis Significance Testing (NHST) and demonstrate the dangers of using significance thresholds to interpret data.

4) Statistical precision is not the same as measurement precision. This will bring attention to the many different types of uncertainties that are built into archaeological datasets (e.g., statistical precision, instrument measurement error, natural variation),.Recognising this is key for drawing reliable inferences from our data.

5) Meta-analyses and forest plots can be useful for synthesising previous research. This will help spread awareness about the benefit of meta-analyses for creating evidence-driven summaries of previous findings.

The discussion draws on examples from isotope archaeologybioarchaeology, and organic residue analysis to illustrate how switching from a reliance on significance testing to a reliance on effect sizes can improve methodological rigour and the representativeness of our findings. The paper ends with a discussion of the roles and responsibilities of supervisors for creating an effective learning environment for statistical training. This includes, but is not limited to, acknowledging the problems of NHST and advocating for adherence to Open Science principles. Ultimately, the changes suggested in this paper will help us raise discipline-wide standards for quantitative training and improve both the breadth and the depth of archaeological research.

Thankyou Jiangang! A Ten-Year Journey to Significance Roulette

Jiangang Xia is an enterprising professor at the University of Nebraska who, among many other things, teaches into China. He alerted me some years ago to the difficulty his students in China had accessing my videos because YouTube was blocked for them. So I mounted the videos at this OSF site.

As another enterprising step, Jiangang has recently been making a number of posts to LinkedIn. Below is one of these:

From a LinkedIn post by Jiangang Xia

Rethinking Quantitative Reasoning in Educational Research (4)

Jiangang Xia

Jiangang Xia

Associate Professor of Educational Administration at University of Nebraska–Lincoln

January 13, 2026

Geoff Cumming and the New Statistics: Estimation as a Way of Thinking

Some years only reveal their importance in hindsight. For me—and for statistics reform—2014 was one of those years.

That was the year I had just begun my academic career at the University of Nebraska–Lincoln. It was also the year Rex Kline visited UNL and delivered his keynote, “Hello, Statistics Reform.” (for Kline’s talk, please see my Post 2 for more details). At the time, I had no idea how much that visit would eventually shape my thinking, teaching, and research.

What I did not know then—and only came to appreciate much later—is that 2014 was also the year Geoff Cumming delivered his workshop “The New Statistics: Estimation and Research Integrity” at the APS Annual Convention in San Francisco.

I was completely unaware of that workshop.

In fact, I would remain largely unaware of Cumming’s work for several more years—even though it quietly passed through my academic life more than once.

Missed encounters (seen only in hindsight)

While preparing this post, I did something simple: I searched Geoff Cumming’s name in my old emails. That’s when I realized how often his work had crossed my path without fully registering.

In 2016, I co-chaired a dissertation in which the student cited Cumming’s influential article “The New Statistics: Why and How” (2014). At the time, neither of us fully grasped what the article was really asking us to reconsider. Although reform language appeared, the dissertation still relied on phrases like “marginally significant” for p values above .05—an indication that NHST logic remained firmly in place. Estimation had been encountered, but not yet learned as a way of thinking.

In 2017, I received an email from Routledge inviting faculty to request a desk copy of Introduction to the New Statistics. I didn’t request it. Another missed opportunity—one I only recognize now, looking backward.

These weren’t personal oversights so much as reflections of how deeply NHST was normalized in our training. Reform ideas were present, but the infrastructure for learning them—how to teach them and how to use them—was still thin.

From critique to alternative

It wasn’t until 2019, through Rex Kline’s work, that the larger picture finally came into focus for me. From there, I traced the reform movement backward—and Geoff Cumming’s role became unmistakable.

Cumming and his colleagues were not simply extending Jacob Cohen’s critique of null hypothesis significance testing. Cohen had already shown, powerfully, why NHST was flawed and had pointed toward alternatives. He also played a central role in the APA Task Force on Statistical Inference, whose report was released in 1999.

Unfortunately, Cohen passed away in 1998, before that work could be fully carried forward. When the APA’s 5th edition was published in 2001, many of the reform ideas were only partially adopted, and everyday research norms remained largely unchanged.

What Cumming did next was different.

Through decades of scholarship—culminating in The New Statistics—he articulated a coherent alternative centered on estimation rather than binary decisions, emphasizing effect sizes, confidence intervals, precision, uncertainty, and cumulative evidence. By the New Statistics, Cumming meant an estimation-centered approach to inference—focusing on effect sizes, confidence intervals, and uncertainty, and on combining evidence across studies through meta-analysis—rather than making binary decisions based on p values alone.

From alternative to institutionalization and teaching

That work did not remain theoretical.

In 2008, Cumming was invited to join the small working group responsible for statistical reporting standards in the APA’s 6th edition, where he was a driving force behind the requirement that effect sizes and confidence intervals be reported for every research question—helping move estimation from an optional supplement to a core reporting standard.

Just as importantly, Cumming took the initiative to teach this alternative—by writing new articles (2014), new textbooks (2013, 2017, 2024), offering workshops (2014 APS), and developing demonstrations (ESCI) aimed at broad audiences, not just methodologists.

Why experience matters in teaching reform

Along the way, Cumming recognized something more fundamental: logical arguments alone were not enough.

As he later reflected, perhaps NHST had become “the researcher’s heroin—an addiction impervious to reason.”

If that was true, persuasion would require more than explanation. It would require experience.

This insight shaped his teaching. Rather than debating p values in the abstract, Cumming showed researchers—often viscerally—how unstable they are through demonstrations such as the dance of the p values, p intervals, and significance roulette. The goal was not just to convince the mind, but to engage the gut—to help researchers feel uncertainty rather than deny it.

A decade later, the field itself began to catch up. In 2024, Cumming’s 2014 article “The New Statistics: Why and How” received SAGE’s 10-Year Impact Award, recognizing research whose influence endures well beyond the standard citation window. The award was a reminder that reform ideas are often understood slowly—resisted early, adopted unevenly, and acknowledged only after they have quietly reshaped teaching and research norms.

Full circle: from missed encounters to transformed practices

What changed for me after 2019 was not just what I read—but how I taught, mentored, and conducted research.

In 2022, I shared Cumming’s 2014 article with my doctoral advisee Amanda. Her dissertation became the first I supervised to fully abandon significance language, adopting estimation-based interpretation throughout. That same year, a manuscript my student Cailen and I submitted—published in 2023—explicitly drew on Cohen, Kline, Cumming, and the ASA (2016) statement. It was my first research article grounded fully in estimation thinking.

And in a quiet but meaningful full circle, when I needed materials for my 2024 summer teaching in China, Geoff Cumming himself shared his 2014 APS workshop videos with me—materials I have since used both internationally and in my quantitative methods courses at UNL.

Looking back now, the reform was happening all around me in 2014.

I just wasn’t ready to see it.

Perhaps that is how intellectual change often works—not as a single revelation, but as a series of missed encounters that eventually align. And perhaps genuine reform requires not only better arguments, but better ways of helping researchers experience uncertainty.

For readers who want to explore further

(All materials shared with permission; enormous credit to Geoff Cumming, Bradley Dean, and Robert Calin-Jageman.)

Dear colleagues,

  • When did you first encounter Geoff Cumming’s work—if at all?
  • Have you seen estimation treated as an add-on, rather than a way of thinking?
  • What important ideas did you meet early, but only understand much later?

#Cumming #NewStatistics #EstimationThinking #QuantitativeMethods #EducationalResearch

Thank you Jiangang!

Geoff

A Statistics Textbook for the AI Era

Miodrag Lovrić is an enormously energetic statistician and educator. He persuaded 700 scholars from 110 countries to contribute to the massive four-volume second edition of the International Encyclopedia of Statistical Science (Springer, 2025).

Now he is close to completing Statistical Thinking for the AI Era, which comprises Vol. 1: Foundations and Inference, and Vol. 2: Advanced Methods–each with about 15 chapters and 520 pages. The chapters I have seen indicate it’s a terrific textbook.

Miodrag has lived and worked on at least four continents and takes a wholeheartedly global focus in both those major works. The textbook starts with Forewords from many continents. Miodrag generously invited me to contribute a foreword from Australia–which appears below. As you see, I slipped in mention that I’m an Antarctic tragic, even if it’s a big claim that I speak also from that seventh continent!

Foreword From Australia and Antarctica

I’ve been an Antarctic tragic since I was a boy. I’ve participated in citizen science during two voyages south. I’ve observed glacial retreat and penguin colonies dying. I read about giant datasets analysed by AI-assisted statistical modelling of changes to sea ice, weather, and much else including terrifying tipping points that will determine catastrophic changes our children must endure.

I’ve given a well-received statistics talk to researchers at the Institute for Marine and Antarctic Studies in a building on the wharf where Antarctic supply ships berth. All around us were whale skulls, ancient sledges, and other memorabilia. I feel I can claim at least a little justification for offering this piece on behalf of the seventh continent, as well as speaking from Australia.

Statistics is about communication and therefore the province of psychologists. Researchers publish representations of their results and readers, ranging from researchers to politicians, policymakers, and ordinary people, draw conclusions from those representations. However, vast industries misrepresent research results as they seek to persuade us to support autocrats, eat unhealthy food, buy things we don’t want that trash the planet, and burn fossil fuels to destroy the chance our children will inherit a liveable world.

My students and I studied people’s misconceptions of basic statistical concepts. That was statistical cognition, which has now developed hugely into metascience—a wonderfully interdisciplinary field that investigates how research is done, and should be done to be more trustworthy. For example, it studies how paper mills use AI tools to generate fake manuscripts that desperate researchers buy then submit to journals. Quality journals use AI tools and much editorial expertise to try to weed out fakes before these criminally pollute the research literature—much of it on life-and-death issues.

Around 2014 Open Science emerged: replication is central; this requires meta-analysis to synthesise results; this requires point and interval estimates, and statistical significance is irrelevant and often damaging. My statistics textbooks advocate estimation and meta-analysis. The intro book, with Bob Calin-Jageman, integrates estimation and Open Science all through. The second edition (Routledge, 2024) includes Bob’s wonderful open source software that’s ideal for beginners and researchers aiming for good estimation and Open Science practices. At thenewstatistics.com is more information, also the significance roulette simulation, which dramatizes a highly misleading feature of p values that few researchers appreciate. The good news is that these new approaches are found accessible by students and are a delight to teach.

I conclude that everyone should have a basic understanding of evidence and statistics. This book is broad in scope, well-informed, and future-focussed. For example, it explains traditional, Bayesian, and randomisation frameworks for estimation, any of which can support Open Science. As I read, I’m often applauding the approaches Professor Lovrić has chosen. With this book he is making an enormous contribution to statistics education around the world.

Geoff

The p Value Casino Is Open–For Significance Roulette!

Excel 2003, the best version ever, was enshittified by MicroSoft in the 2007 version, which was way slower and dropped many wonderful animation facilities :-(. Even vast efforts would not get my great Significance Roulette simulation running in the new version.

Two videos use my Excel 2003 version to explain: https://tiny.cc/SigRoulette1 and https://tiny.cc/SigRoulette2 or simply search at YouTube for ‘significance roulette’.

Now Bradley Dean has built Significance Roulette to run in your browser–as part of esci web.

Significance Roulette, after an initial study gives p = .01

It has taken me close to 20 years to build the following argument that leads to significance roulette, and which is summarised in the first half of our open access article Calin-Jageman & Cumming, 2024

  • For approaching a century numerous distinguished scholars, including philosophers of science, statisticians, and psychologists, have published cogent critiques of p values, significance testing, and how researchers across science use these.
  • Even so, a large proportion of researchers, teachers, journals, and granting bodies use null hypothesis significance testing (NHST)–often based on p < .05 or p < .01–as the standard for concluding whether or not an effect exists, whether or not a result is large or important. Despite this logic being wrong-headed in so many ways!
  • Devotion to NHST and p <.05 resembles an addiction–the researcher’s heroin. Rational argument is not sufficient to shake the addiction. Could a dramatic demonstration, perhaps persuading via the gut rather than the brain, shake this addiction?
  • A striking but little-known feature of p values is that they are highly unreliable–their sampling variability is astonishingly large. Replicate a study, exactly the same but with a new random sample, and expect to obtain replication p that can take just about any value between 0 and 1!
  • Jerry Lai in his lovely PhD studies took three converging approaches to find that a large proportion of published researchers in psychology, medicine, and statistics severely under-estimate the amount of variability in the p value with replication.
  • My first demonstration of p value variability was the dance of the p values, the first video of which dates from 2009. Search YouTube for ‘dance of the p values’ to find several videos. You can also play with the dances in esci web.
  • My second approach arose from study of the probability distribution of replication p, the p value obtained in a replication. My highly-cited 2008 article has details.
  • I define the p interval as the 80% prediction interval for replication p. It’s astonishingly long! For example, if an initial study obtains p = .05, the p interval is (.0002, .65), meaning an 80% chance of p within that interval and fully a 10% chance it falls below .0002 and 10% above .65. After p = .01 the interval is (.00001, .41). After initial p = .001 (*** highly statistically significant) there is fully a 1 in 6 chance a replication does not even achieve p < .05! After p = .20 ns there is a 1 in 3 chance a replication finds p < .05!
  • I take the probability distribution of replication p following initial p = .01 and divide the area under the curve, which represents probability, into 38 equal areas. I label each area with the p value in the centre of the interval. I have 38 p values, a few large and many small and very small, which accurately represent that probability distribution.
  • I scatter those 38 p values randomly around the 38 bins of a roulette wheel. Simply click to spin, wait for a few moments, and see the replication p you might have obtained from a replication. Much faster and cheaper than the hassle of raising a grant, recruiting participants, hassling with ethics approval… and collecting data!
  • At Significance Roulette in esci web you can click between initial p of .05 and .01. With initial p = .01, replication p values tend a little smaller, of course. But the striking thing is how widely spread the p values are! The distributions are pictured to left of the wheel. For example, switch from .05 to .01 and note a slightly smaller number of deep blue, deep trombone sound, despairing figures for p > .1 ns and slightly more bright red, triumphant trumpet blast, elated figures for p < .001 ***. Click SPIN, and note your quickened heart beat, sweaty palms, and that you are holding your breath–will you be despairing or elated–and in only a few seconds you’ll know!

Will this approach to tackling addiction via the gut be more effective than the decades of argument addressed at the cortex?

Best of luck at the p Value Casino!

Take-home messages

  • p values are unbelievably unreliable
  • Any p value could easily have been just about any other value
  • No p value deserves our trust 🙁
  • Simply don’t use p values, there are much better ways 🙂

Geoff

P.S. Enormous thanks to Bradley and Bob, who made it all happen.

“The New Statistics” by INXS: Why Not?!

Thanks to Prof Dena A. Pastor, of James Madison University for this suggestion.

Whenever you hear the hit New Sensation, by Australian rock band INXS, replace “A new sensation” with “The New Statistics”. This works for me (of course it would), even if the ’80s are a bit recent for me.

Enjoy!

Geoff

The ideal statistics curriculum?

This week I (Bob) was part of a workshop on the future of neuroscience education. I got to be part of a rockstar panel with Tari Tan (Harvard), Monica Linden (Brown), and Rosalind Segal (Harvard). The event was organized by Lique Coolen (Kent State).

Tari spoke about curricular goals for neuroscience, reviewing an SFN initiative to define core competencies for trainees at the undergraduate and graduate level. I didn’t know about this! Tari gave a great overview; it’s well worth checking out: https://www.sfn.org/careers/higher-education-and-training/core-competencies

Monica discussed the role of generative AI in the future of neuroscience education and gave lots of thoughtful examples and resources.

Rosalind discussed developing an internship program for the neuroscience *PhD* program at Harvard! Very thougtful and interesting way for grad students to explore the diverse career paths that can come out of doctoral training — but also raised lots of interesting issues (a stated goal of internships that don’t impede research seemed hard to balance; students seemed to express worries about losing faculty support if they expressed ‘wrong’ career goals… lots to think about!).

My talk was about the ideal statistics curriculum. My short answer was “There isn’t one!” — neuro is too diverse and there is not only one ‘right’ way to do statistical inference. Still, the field of neuroscience shows some evident difficulties making valid claims from data (on that note, check out Chen et al., 2024), and so at least thinking about some broad ideals for training seems like a useful exercise. I came up with these guiding principles:

  • Estimation thinking at the forefront
  • Teach using simulations — for exploration at first, but eventually for planning experiments before they are conducted
  • Quickly graduate from toy examples to real, complex data sets and projects that require critical thinking about Multiplicity and Interdependence, and teach robustness checking to help students validate their approaches.
  • Integrate Open Science throughout
  • Ensure training is not just about statistics, but about the ‘neglected factors’, including good design and measurement — strength of evidence and quality of research are about so much more than just the statistics generated, and improvements are about so much more than sample size.

My talk and some resources on each of these topics are here: https://osf.io/muy6u/wiki/Workshops%20-%20SFN%202024/

We’re collating all the resources from all the speakers; I’ll post a link when I have that.

1635500 {1635500:JY7HWM5Q},{1635500:QX7U9XKT} 1 apa 50 default 2638 https://thenewstatistics.com/itns/wp-content/plugins/zotpress/
%7B%22status%22%3A%22success%22%2C%22updateneeded%22%3Afalse%2C%22instance%22%3Afalse%2C%22meta%22%3A%7B%22request_last%22%3A0%2C%22request_next%22%3A0%2C%22used_cache%22%3Atrue%7D%2C%22data%22%3A%5B%7B%22key%22%3A%22JY7HWM5Q%22%2C%22library%22%3A%7B%22id%22%3A1635500%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Chen%22%2C%22parsedDate%22%3A%222022-05-01%22%2C%22numChildren%22%3A2%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BChen%2C%20D.%20%282022%29.%20%26lt%3Bi%26gt%3BSTATISTICAL%20PRACTICE%20IN%20PRECLINICAL%20NEUROSCIENCES%3A%20IMPLICATIONS%20FOR%20SUCCESSFUL%20TRANSLATION%20OF%20RESEARCH%20EVIDENCE%20FROM%20ANIMALS%20TO%20HUMANS%20Committee%20Member%26lt%3B%5C%2Fi%26gt%3B.%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22book%22%2C%22title%22%3A%22STATISTICAL%20PRACTICE%20IN%20PRECLINICAL%20NEUROSCIENCES%3A%20IMPLICATIONS%20FOR%20SUCCESSFUL%20TRANSLATION%20OF%20RESEARCH%20EVIDENCE%20FROM%20ANIMALS%20TO%20HUMANS%20Committee%20Member%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Darin%22%2C%22lastName%22%3A%22Chen%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22date%22%3A%222022-05-01%22%2C%22originalDate%22%3A%22%22%2C%22originalPublisher%22%3A%22%22%2C%22originalPlace%22%3A%22%22%2C%22format%22%3A%22%22%2C%22ISBN%22%3A%22%22%2C%22DOI%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222024-10-08T15%3A15%3A56Z%22%7D%7D%2C%7B%22key%22%3A%22QX7U9XKT%22%2C%22library%22%3A%7B%22id%22%3A1635500%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Calin-Jageman%20and%20Cumming%22%2C%22parsedDate%22%3A%222019%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BCalin-Jageman%2C%20R.%20J.%2C%20%26amp%3B%20Cumming%2C%20G.%20%282019%29.%20Estimation%20for%20Better%20Inference%20in%20Neuroscience.%20%26lt%3Bi%26gt%3BEneuro%26lt%3B%5C%2Fi%26gt%3B%2C%20%26lt%3Bi%26gt%3B6%26lt%3B%5C%2Fi%26gt%3B%284%29%2C%20ENEURO.0205-19.2019.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.1523%5C%2FENEURO.0205-19.2019%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.1523%5C%2FENEURO.0205-19.2019%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22journalArticle%22%2C%22title%22%3A%22Estimation%20for%20Better%20Inference%20in%20Neuroscience%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Robert%20J.%22%2C%22lastName%22%3A%22Calin-Jageman%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Geoff%22%2C%22lastName%22%3A%22Cumming%22%7D%5D%2C%22abstractNote%22%3A%22The%20estimation%20approach%20to%20inference%20emphasizes%20reporting%20effect%20sizes%20with%20expressions%20of%20uncertainty%20%28interval%20estimates%29.%20In%20this%20perspective%20we%20explain%20the%20estimation%20approach%20and%20describe%20how%20it%20can%20help%20nudge%20neuroscientists%20toward%20a%20more%20productive%20research%20cycle%20by%20fostering%20better%20planning%2C%20more%20thoughtful%20interpretation%2C%20and%20more%20balanced%20evaluation%20of%20evidence.%22%2C%22date%22%3A%2207%5C%2F2019%22%2C%22section%22%3A%22%22%2C%22partNumber%22%3A%22%22%2C%22partTitle%22%3A%22%22%2C%22DOI%22%3A%2210.1523%5C%2FENEURO.0205-19.2019%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Fwww.eneuro.org%5C%2Flookup%5C%2Fdoi%5C%2F10.1523%5C%2FENEURO.0205-19.2019%22%2C%22PMID%22%3A%22%22%2C%22PMCID%22%3A%22%22%2C%22ISSN%22%3A%222373-2822%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%226ECZTSSQ%22%2C%22M4C95AAC%22%5D%2C%22dateModified%22%3A%222023-01-17T22%3A06%3A18Z%22%7D%7D%5D%7D
Chen, D. (2022). STATISTICAL PRACTICE IN PRECLINICAL NEUROSCIENCES: IMPLICATIONS FOR SUCCESSFUL TRANSLATION OF RESEARCH EVIDENCE FROM ANIMALS TO HUMANS Committee Member.
Calin-Jageman, R. J., & Cumming, G. (2019). Estimation for Better Inference in Neuroscience. Eneuro, 6(4), ENEURO.0205-19.2019. https://doi.org/10.1523/ENEURO.0205-19.2019

Ten Years On, The New Statistics Workshop Videos Still Live

2014 was exciting: In January, Psychological Science‘s new requirements introduced many psychologists to Open Science. In May the APS Convention in San Francisco was an exhilarating (for me, anyway) whirl of symposia, talks, and discussions about how Open Science should be shaped and advanced. APS made 6 videos of my workshop ‘The New Statistics: Estimation and Research Integrity’. (Or use tiny.cc/apstnsvideos)

“Don’t we LOVE our p values!” I tossed the large p away, then jumped on it. Alas, my jumping didn’t make it into any of the six videos.

I was recently told the link to my workshop .ppt slide files was dead. I updated it, so tiny.cc/geoffdocs now points to a new OSF page with those files.

That led me to revisit the videos. They seem to me to hold up well–maybe even still worth considering. My one suggestion is to change every ‘research integrity‘ into ‘Open Science‘. OS is a way better term and what I would have used even a few months later.

It was the big Psychological Science changes and my article The New Statistics: Why and How that led to the workshop and videos. (For the article: tiny.cc/tnswhyhow)

With the recent Sage award for that article, it’s maybe timely to revisit the videos and again make the .ppt files accessible.

Here’s a list of the videos:

Part 1: Confidence Intervals, NHST, and p Values

Part 2: Research Integrity and the New Statistics

Part 3: Effect Sizes and Confidence Intervals

Part 4: The New Statistics in Action

Part 5: Planning, Power and Precision

Part 6: Meta-analysis and Meta-analytic thinking

Geoff

‘The New Statistics’ (2013) Wins Sage 10-Year Impact Award

The New Statistics: Why and How (abstract below) explained the advantages of moving on from NHST to the new statistics (estimation and meta-analysis) and the need for better practices to improve research integrity. I’m delighted that an award from Sage indicates the article seems to be helping researchers improve what they do. Next: Can ITNS2 help the next generation do even better?

The article appeared online in late 2013, so was considered when Sage examined the citation numbers of all articles appearing in any of the 400+ journals Sage published back in 2013. It was one of the top three most cited, so has been given a Sage 10-Year Impact Award. Sage’s announcement is here. Sage has just published a blog post about it with headline:

Now for the abstract:

The article was commissioned by Eric Eich, then editor-in-chief of Psychological Science, to appear immediately following his famous editorial Business Not As Usual. This opened the Journal’s first issue of 2014 and announced sweeping changes in the journal’s submission requirements, which, for many psychologists, marked the arrival of Open Science.

Interview

Sage’s blog post includes an email interview with me. Here’s a brief summary:

What was it in your own background that led to your article?

When I was a teenager my father gave me a simple explanation of significance testing. I said something like “That’s weird, sort of backwards. And why .05?” He replied “I agree, but that’s the way we do it.”

Over decades of teaching I became ever more dissatisfied with NHST, and focused ever more on confidence intervals (CIs).

Was there an article that had a particularly strong influence on you?

Frank Schmidt (1996) wrote: “It is now possible to use meta-analysis to show that reliance on significance testing retards the development of cumulative knowledge.” A revelation!

What did Schmidt’s article lead to?

About 2003 I started using an Excel forest plot to give a simple explanation of meta-analysis in my intro course. I was delighted: Students told me it just made sense. Of course, for meta-analysis you need a CI from each study, while p values are irrelevant, even misleading.

      Figure: Dances of means, confidence intervals, and p values.

In 2009 I uploaded a video of the dance of the p values. I became passionate about advocating the new statistics (estimation and meta-analysis). I wrote Understanding The New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis (UTNS, 2012).

What was happening in psychology at about that time?

Ioannidis (2005) explained how reliance on NHST was a major cause of the replication crisis. Largely in response to that crisis, Open Science arrived—perhaps the most important advance in how science is done for a very long time.

Eric Eich’s famous editorial Business Not As Usual in the January 2014 issue of Psychological Science marked the arrival of Open Science in psychology. Months earlier Eric had invited me to write a tutorial article to support the changes he wanted. This was The New Statistics: Why and How and was published immediately following his editorial.

What has been the reception of the article?

Mainly very positive. Some have felt I went too far in advising that in most cases it’s better not to use NHST at all. Some Bayesians have been unhappy with the focus on confidence intervals.

Revisiting that article, what would you have done differently?

I used the term ‘research integrity’, but ‘Open Science’ was coming into use and I soon realized that was way better. Reading the article today, for ‘research integrity’ read ‘Open Science’.

Otherwise, I think the article has held up well, including all 25 guidelines in Table 1.

What has happened since

Psychological Science has continued to lead in the adoption of Open Science practices.

Meta-science, also known as meta-research, has emerged and now thrives as a highly multi-disciplinary field. It applies the scientific method to improve that method—wonderful!

What have you been doing since?

I teamed with Robert Calin-Jageman to write the first intro statistics textbook based on the new statistics and with Open Science all through. The second edition has just come out: Introduction to The New Statistics: Estimation, Open Science, and Beyond, 2nd edition (ITNS2, 2024). It has much improved software, as we explain in Calin-Jageman & Cumming (2024), which is on open access.

We believe this book can sweep the world—we’ll see! To read the Preface and Chapter 1 go to www.thenewstatistics.com. In the second para is a link to the book’s Amazon site. Click ‘Read sample’.  

References

Calin-Jageman, R., & Geoff Cumming, G. (2024). From significance testing to estimation and Open Science: How esci can help. International Journal of Psychology,     https://doi.org/10.1002/ijop.13132

Cumming, G. (2012). The New Statistics: Effect sizes, confidence intervals, and meta-analysis. New York: Routledge. 

Cumming, G. (2014) The new statistics: Why and how. Psychological Science. 25(1), 7-29. https://doi.org/10.1177/0956797613504966

Cumming, G., & Calin-Jageman, R. (2024). Introduction to The New Statistics: Estimation, Open Science, & Beyond, 2nd edition. New York: Routledge.

Eich, E. (2014) Business not as usual. Psychological Science, 25(1), 3–6. https://doi.org/10.1177/0956797613512465

Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine 2: e124. https://doi.org/10.1371/journal.pmed.0020124

Schmidt, F. L. (1996). Statistical significance testing and cumulative knowledge in psychology: Implications for training of researchers. Psychological Methods, 1(2), 115-129. https://doi.org./10.1037/1082-989X.1.2.115 

Booklisti. For Finding Interesting Books, Now Including ITNS2

Exploring Booklisti is a neat way to find good books to read.

Booklisti, would you believe, comprises lots of short lists of books that hang together. I have two lists, ITNS2 appearing in each. My first list is just UTNS, my first book, and ITNS2, our second edition of the intro book.

My second list (go to that site for links to the books listed below), titled Open Science and how to do better research with better statistics, comprises:

Introduction to The New Statistics, Second Edition By Geoff Cumming and Robert Calin-Jageman

Understanding The New Statistics By Geoff Cumming

A Student’s Guide to Open Science By Charlotte Pennington

The Seven Deadly Sins of Psychology: A Manifesto for Reforming the Culture of Scientific Practice By Chris Chambers

Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth By Stuart Ritchie

Beyond Significance Testing By Rex B. Kline

Research Methods in Psychology: Evaluating a World of Information By Beth Morling

The Design of Experiments in Neuroscience By Mary E. Harrington

Happy exploring of Booklisti and happy reading.

Geoff

Estimation, Open Science, and Bob’s Wonderful New esci

Our open access article just released at https://doi.org/10.1002/ijop.13132:

Highlights

  • Three dramatisations of the enormous unreliability of the p value. Can these help weaken researchers’ addiction to NHST that has withstood more than half a century of cogent rational critiques?
  • Bob’s wonderful new open-source esci software with great estimation-based figures: See worked examples, and work along if you wish.

Abstract

We argue that researchers should test less, estimate more, and adopt Open Science practices. We outline some of the flaws of null hypothesis significance testing and take three approaches to demonstrating the unreliability of the p value. We explain some advantages of estimation and meta-analysis (“the new statistics”), especially as contributions to Open Science practices, which aim to increase the openness, integrity, and replicability of research. We then describe esci (estimation statistics with confidence intervals): a set of online simulations, and an R package for estimation that integrates into jamovi and JASP. This software provides (a) online activities to sharpen understanding of statistical concepts (e.g., “The Dance of the Means”); (b) effects sizes and confidence intervals for a range of study designs, largely by using techniques recently developed by Bonett; (c) publication-ready visualisations that make uncertainty salient; and (d) the option to conduct strong, fair hypothesis evaluation through specification of an interval null. Although developed specifically to support undergraduate learning through the 2nd edition of our textbook, esci should prove a valuable tool for graduate students and researchers interested in adopting the estimation approach. Further information is at https://thenewstatistics.com.

Figure 1. Significance roulette. If an initial study obtains p=.01, an exact replication–just the same but with an new sample–will obtain a p value drawn from the enormous spread of values on the wheel.

The Enormous Unreliability of p

This is the first time (1) the dance of the p values (search YouTube), (2) significance roulette (Figure 1; and search YouTube), and (3) p intervals (see the article) have all been described together in print. Significance roulette has been around for a while but this is its first outing in print. Alas, p values simply don’t deserve our trust. Enjoy the figures!

esci web

This component of esci is a set of simulations and tools by our colleague Gordon Moore that run in any browser. Explore the dances, play with sampling distributions, find critical values, and more.

esci for Data Analysis

Bob’s esci is an open-source package in R, which can be run in R, or within jamovi or (by December 2024) in JASP. We describe the wide range of measures and designs esci can analyse, including meta-analysis, and work through several examples. We emphasise figures that highlight uncertainty, especially by picturing confidence intervals.

Figure 2. Part of jamovi screen showing selection of the ‘Gender math IAT’ data file for opening.

We argue that p values, if used at all, are most valuable in the context of hypothesis evaluation based on an interval null hypothesis, and best understood with the help of an esci figure–see Figure 3.

Interactions can be challenging to understand and interpret; again esci provides figures designed to help–see Figure 4.

Pro Tip: Data Files Now in esci

The article advises download of jamovi-format data files (Gender math IAT.omvGender math IAT ma.omvCampus Involvement.omv and MeditationBrain.omv) from https://osf.io/uhwj2. Since the final version of the article was submitted Bob has integrated into esci all the data files used in ITNS2, including these four, so download from OSF is no longer needed.

Figure 3. esci figure for two independent groups. At left, the data points, means and 95% CIs for the two groups. The black triangle marks the difference between the means. This and its 90% and 95% CIs are shown on the difference axis at right. The two CIs allow test of the interval null hypothesis, the pink stripe.

Examples of Analyses by esci

Figures 3 and 4 are just two illustrations from the example esci analyses discussed in the article.

To open a data file within esci, click top left in jamovi, then click Open, Data Library, and scroll to see all the data files for ITNS2 arranged by chapter. Figure 2 shows selection of the first example file used in the article.

Figure 3 is an esci figure from a two independent groups analysis of the Gender math IAT file. The grey areas on the CIs are what we call plausibility curves. These illustrate variation in the plausibility, or relative likelihood, that values across and beyond the interval are the population value.

Figure 4 is one of the ways esci can display a 2 x 2 interaction–part of an RCT analysis of the MeditationBrain file.

If you wish, work along with the examples. The rich UI (user interface) of esci gives lots of scope to make figures look just as you want them–there’s advice about how to tweak your figures to look like those in the article.

As ever, we’d love to hear your comments on the new book and new software. Enjoy.

Figure 4. One way esci displays a 2 x 2 interaction. The difference in slope of the two lines indicates the size of the interaction. The fans of faint lines give a rough indication of the extent of uncertainty in estimating the slopes of the lines.

Geoff