‘The New Statistics’ (2013) Wins Sage 10-Year Impact Award

The New Statistics: Why and How (abstract below) explained the advantages of moving on from NHST to the new statistics (estimation and meta-analysis) and the need for better practices to improve research integrity. I’m delighted that an award from Sage indicates the article seems to be helping researchers improve what they do. Next: Can ITNS2 help the next generation do even better?

The article appeared online in late 2013, so was considered when Sage examined the citation numbers of all articles appearing in any of the 400+ journals Sage published back in 2013. It was one of the top three most cited, so has been given a Sage 10-Year Impact Award. Sage’s announcement is here. Sage has just published a blog post about it with headline:

Now for the abstract:

The article was commissioned by Eric Eich, then editor-in-chief of Psychological Science, to appear immediately following his famous editorial Business Not As Usual. This opened the Journal’s first issue of 2014 and announced sweeping changes in the journal’s submission requirements, which, for many psychologists, marked the arrival of Open Science.

Interview

Sage’s blog post includes an email interview with me. Here’s a brief summary:

What was it in your own background that led to your article?

When I was a teenager my father gave me a simple explanation of significance testing. I said something like “That’s weird, sort of backwards. And why .05?” He replied “I agree, but that’s the way we do it.”

Over decades of teaching I became ever more dissatisfied with NHST, and focused ever more on confidence intervals (CIs).

Was there an article that had a particularly strong influence on you?

Frank Schmidt (1996) wrote: “It is now possible to use meta-analysis to show that reliance on significance testing retards the development of cumulative knowledge.” A revelation!

What did Schmidt’s article lead to?

About 2003 I started using an Excel forest plot to give a simple explanation of meta-analysis in my intro course. I was delighted: Students told me it just made sense. Of course, for meta-analysis you need a CI from each study, while p values are irrelevant, even misleading.

      Figure: Dances of means, confidence intervals, and p values.

In 2009 I uploaded a video of the dance of the p values. I became passionate about advocating the new statistics (estimation and meta-analysis). I wrote Understanding The New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis (UTNS, 2012).

What was happening in psychology at about that time?

Ioannidis (2005) explained how reliance on NHST was a major cause of the replication crisis. Largely in response to that crisis, Open Science arrived—perhaps the most important advance in how science is done for a very long time.

Eric Eich’s famous editorial Business Not As Usual in the January 2014 issue of Psychological Science marked the arrival of Open Science in psychology. Months earlier Eric had invited me to write a tutorial article to support the changes he wanted. This was The New Statistics: Why and How and was published immediately following his editorial.

What has been the reception of the article?

Mainly very positive. Some have felt I went too far in advising that in most cases it’s better not to use NHST at all. Some Bayesians have been unhappy with the focus on confidence intervals.

Revisiting that article, what would you have done differently?

I used the term ‘research integrity’, but ‘Open Science’ was coming into use and I soon realized that was way better. Reading the article today, for ‘research integrity’ read ‘Open Science’.

Otherwise, I think the article has held up well, including all 25 guidelines in Table 1.

What has happened since

Psychological Science has continued to lead in the adoption of Open Science practices.

Meta-science, also known as meta-research, has emerged and now thrives as a highly multi-disciplinary field. It applies the scientific method to improve that method—wonderful!

What have you been doing since?

I teamed with Robert Calin-Jageman to write the first intro statistics textbook based on the new statistics and with Open Science all through. The second edition has just come out: Introduction to The New Statistics: Estimation, Open Science, and Beyond, 2nd edition (ITNS2, 2024). It has much improved software, as we explain in Calin-Jageman & Cumming (2024), which is on open access.

We believe this book can sweep the world—we’ll see! To read the Preface and Chapter 1 go to www.thenewstatistics.com. In the second para is a link to the book’s Amazon site. Click ‘Read sample’.  

References

Calin-Jageman, R., & Geoff Cumming, G. (2024). From significance testing to estimation and Open Science: How esci can help. International Journal of Psychology,     https://doi.org/10.1002/ijop.13132

Cumming, G. (2012). The New Statistics: Effect sizes, confidence intervals, and meta-analysis. New York: Routledge. 

Cumming, G. (2014) The new statistics: Why and how. Psychological Science. 25(1), 7-29. https://doi.org/10.1177/0956797613504966

Cumming, G., & Calin-Jageman, R. (2024). Introduction to The New Statistics: Estimation, Open Science, & Beyond, 2nd edition. New York: Routledge.

Eich, E. (2014) Business not as usual. Psychological Science, 25(1), 3–6. https://doi.org/10.1177/0956797613512465

Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine 2: e124. https://doi.org/10.1371/journal.pmed.0020124

Schmidt, F. L. (1996). Statistical significance testing and cumulative knowledge in psychology: Implications for training of researchers. Psychological Methods, 1(2), 115-129. https://doi.org./10.1037/1082-989X.1.2.115 

Booklisti. For Finding Interesting Books, Now Including ITNS2

Exploring Booklisti is a neat way to find good books to read.

Booklisti, would you believe, comprises lots of short lists of books that hang together. I have two lists, ITNS2 appearing in each. My first list is just UTNS, my first book, and ITNS2, our second edition of the intro book.

My second list (go to that site for links to the books listed below), titled Open Science and how to do better research with better statistics, comprises:

Introduction to The New Statistics, Second Edition By Geoff Cumming and Robert Calin-Jageman

Understanding The New Statistics By Geoff Cumming

A Student’s Guide to Open Science By Charlotte Pennington

The Seven Deadly Sins of Psychology: A Manifesto for Reforming the Culture of Scientific Practice By Chris Chambers

Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth By Stuart Ritchie

Beyond Significance Testing By Rex B. Kline

Research Methods in Psychology: Evaluating a World of Information By Beth Morling

The Design of Experiments in Neuroscience By Mary E. Harrington

Happy exploring of Booklisti and happy reading.

Geoff

Meet Petra: Enthusiasm for Archaeology, Open Science, and Better Statistics

Lunch with Petra Vaiglova

It was a pleasure to meet Petra Vaiglova a few days ago while she was in Melbourne for an archaeological science conference. Fiona Fidler joined us for lunch–thanks to her for hosting.

Originally from the Czech Republic (Czechia), Petra has lived, studied, and worked all over the world, as you can see at her site. Her doctorate is from Oxford. Her TEDx talk outlines some of her research interests.

She arrived in Australia a couple of years ago as a post doc at Griffith University, Queensland. Soon after, she launched into organising what became a three-day online Workshop on Good Statistical Practice in Archaeology–open to anyone of any discipline.

I first learned of her enthusiasm for statistical reform and Open Science when she kindly invited me to speak at that workshop. I gave a talk on the new statistics, and another on using Bob’s new esci software in jamovi–in archaeological science, as in any other discipline. I posted about the workshop here.

Earlier this year Petra took up a lectureship at ANU in Canberra, and enthusiastically took on teaching statistics and topics in archaeological science to both undergraduate and postgraduate students.

She expressed keen interest in our second edition, even volunteering to help. Over many months last year she worked through final drafts of most chapters, picking up many errors and infelicities. Later she worked painstakingly through the proofs of many chapters, picking up elusive tiny errors. She made an immense contribution to ITNS2, as Bob and I acknowledge on p. xxvii.

So last week Petra, Fiona and I had much to discuss. Petra will be at AIMOS in November (more details here), no doubt meeting many of the good folks who make the Meta-science scene in Australia so lively and multi-disciplinary.

As an associate editor of the Journal of Archaeological Science she is helping develop guidelines for that journal to encourage reproducibility and Open Science practices.

I discovered years ago that archaeological science has Open Science lessons for us all–see my post here. I wish Petra all strength for her continuing efforts towards statistical reform and Open Science!

Geoff

Estimation, Open Science, and Bob’s Wonderful New esci

Our open access article just released at https://doi.org/10.1002/ijop.13132:

Highlights

  • Three dramatisations of the enormous unreliability of the p value. Can these help weaken researchers’ addiction to NHST that has withstood more than half a century of cogent rational critiques?
  • Bob’s wonderful new open-source esci software with great estimation-based figures: See worked examples, and work along if you wish.

Abstract

We argue that researchers should test less, estimate more, and adopt Open Science practices. We outline some of the flaws of null hypothesis significance testing and take three approaches to demonstrating the unreliability of the p value. We explain some advantages of estimation and meta-analysis (“the new statistics”), especially as contributions to Open Science practices, which aim to increase the openness, integrity, and replicability of research. We then describe esci (estimation statistics with confidence intervals): a set of online simulations, and an R package for estimation that integrates into jamovi and JASP. This software provides (a) online activities to sharpen understanding of statistical concepts (e.g., “The Dance of the Means”); (b) effects sizes and confidence intervals for a range of study designs, largely by using techniques recently developed by Bonett; (c) publication-ready visualisations that make uncertainty salient; and (d) the option to conduct strong, fair hypothesis evaluation through specification of an interval null. Although developed specifically to support undergraduate learning through the 2nd edition of our textbook, esci should prove a valuable tool for graduate students and researchers interested in adopting the estimation approach. Further information is at https://thenewstatistics.com.

Figure 1. Significance roulette. If an initial study obtains p=.01, an exact replication–just the same but with an new sample–will obtain a p value drawn from the enormous spread of values on the wheel.

The Enormous Unreliability of p

This is the first time (1) the dance of the p values (search YouTube), (2) significance roulette (Figure 1; and search YouTube), and (3) p intervals (see the article) have all been described together in print. Significance roulette has been around for a while but this is its first outing in print. Alas, p values simply don’t deserve our trust. Enjoy the figures!

esci web

This component of esci is a set of simulations and tools by our colleague Gordon Moore that run in any browser. Explore the dances, play with sampling distributions, find critical values, and more.

esci for Data Analysis

Bob’s esci is an open-source package in R, which can be run in R, or within jamovi or (by December 2024) in JASP. We describe the wide range of measures and designs esci can analyse, including meta-analysis, and work through several examples. We emphasise figures that highlight uncertainty, especially by picturing confidence intervals.

Figure 2. Part of jamovi screen showing selection of the ‘Gender math IAT’ data file for opening.

We argue that p values, if used at all, are most valuable in the context of hypothesis evaluation based on an interval null hypothesis, and best understood with the help of an esci figure–see Figure 3.

Interactions can be challenging to understand and interpret; again esci provides figures designed to help–see Figure 4.

Pro Tip: Data Files Now in esci

The article advises download of jamovi-format data files (Gender math IAT.omvGender math IAT ma.omvCampus Involvement.omv and MeditationBrain.omv) from https://osf.io/uhwj2. Since the final version of the article was submitted Bob has integrated into esci all the data files used in ITNS2, including these four, so download from OSF is no longer needed.

Figure 3. esci figure for two independent groups. At left, the data points, means and 95% CIs for the two groups. The black triangle marks the difference between the means. This and its 90% and 95% CIs are shown on the difference axis at right. The two CIs allow test of the interval null hypothesis, the pink stripe.

Examples of Analyses by esci

Figures 3 and 4 are just two illustrations from the example esci analyses discussed in the article.

To open a data file within esci, click top left in jamovi, then click Open, Data Library, and scroll to see all the data files for ITNS2 arranged by chapter. Figure 2 shows selection of the first example file used in the article.

Figure 3 is an esci figure from a two independent groups analysis of the Gender math IAT file. The grey areas on the CIs are what we call plausibility curves. These illustrate variation in the plausibility, or relative likelihood, that values across and beyond the interval are the population value.

Figure 4 is one of the ways esci can display a 2 x 2 interaction–part of an RCT analysis of the MeditationBrain file.

If you wish, work along with the examples. The rich UI (user interface) of esci gives lots of scope to make figures look just as you want them–there’s advice about how to tweak your figures to look like those in the article.

As ever, we’d love to hear your comments on the new book and new software. Enjoy.

Figure 4. One way esci displays a 2 x 2 interaction. The difference in slope of the two lines indicates the size of the interaction. The fans of faint lines give a rough indication of the extent of uncertainty in estimating the slopes of the lines.

Geoff

To Find Interesting Books, Explore bookdna.com, Now Including ITNS2

A couple of years back I posted (here) about bookdna.com, which has gone from strength to strength as an engaging way to browse books for interesting finds. I’ve updated our bookdna entry to ITNS2 and tweaked our recommendations, with Pennington‘s little gem on Open Science now first on the list. The start of our entry:

Our recommended list is now:

A Student’s Guide to Open Science By Charlotte Pennington

Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth By Stuart Ritchie

Beyond Significance Testing By Rex B. Kline

Research Methods in Psychology: Evaluating a World of Information By Beth Morling

The Design of Experiments in Neuroscience By Mary E. Harrington

bookdna offers lots of ways to explore. Happy reading.

Geoff

Vale Bob Rosenthal, Statistical Reform Leader and Much Else

I was much saddened to read of the death last month of Bob Rosenthal. See this obituary; and another in the New York Times.

I met him first in 1996 when I called on him at Harvard to discuss statistical reform. What a gentle, encouraging, and thoroughly nice person! What a giant intellect! He loved nothing better than to find innovative solutions to tricky problems.

Considering statistical reform and Open Science:

  • He was an early proponent of a focus on effect sizes, especially his favourite, Pearson correlation, r.
  • He was a pioneer of meta-analysis and identified what he called the file-drawer effect.
  • Rosenthal and Gaito (1963) reported evidence that researchers’ confidence in an effect drops sharply as the p value increases past .05; they labelled this the cliff effect. This was an early example of statistical cognition–the empirical study of how people understand statistical concepts and reports. We still need much more of that, imho.
  • Around 2009 Jerry Lai wanted to investigate the cliff effect as part of his PhD. He sent Bob a very polite request for any further information about the original study. Promptly, back came an encouraging message to Jerry and a scan of several hand-written pages of the original data. From almost 50 years earlier! A wonderful example of Open Data (well, available data), with no excuses about hard disk crashes and superseded storage formats.
  • He advocated analysis of well-chosen contrasts as better than the customary reliance on Anova and p values (*, **, ***, or ns) to interpret omnibus main and interaction effects. He stated that “the problem is that omnibus tests … do not usually tell us anything we really want to know”. Contrast Analysis: Focused Comparisons in the Analysis of Variance (1985) by Rosenthal and Rosnow remains an accessible and powerful explanation. UTNS, and both editions of ITNS take this planned contrast approach (these days, with preregistration) to the analysis of complex designs.
  • In 2008 Fiona Fidler and I were working on Confidence Intervals : Better Answers to Better Questions. We sent a draft to Bob who was working on an accompanying article Effect Sizes : Why, When, and How to Use Them. Bob responded with enthusiasm, saying he loved our article and also offering valuable suggestions.

Bob’s nickname among his students was “Prof ARRRZZZental“, recognising his love of correlation r.

I salute his memory and his enduring contribution to improving how we do things.

Geoff

Brian Nosek Speaks: A BJKS Podcast

Brian tells his story, and that of the Center for Open Science and the Open Science Framework. A great listen.

Here are the sections:

00:00: Brian’s early interest in improving science
15:24: How the Center for Open Science got funded (by John and Laura Arnold)
26:08: How long is COS financed into the future?
29:01: What if COS isn’t benefitting science anymore?
35:42: Is Brian a scientist or an entrepreneur?
40:58: The future of the Center for Open Science
51:13: A book or paper more people should read
54:42: Something Brian wishes he’d learnt sooner
58:53: Advice for PhD students/postdocs

I recently posted about other BJKS podcastsBenjamin James Kuper-Smith talking with Simine Vazire, Chris Chambers, and me, among others.

Happy listening (or even reading the transcripts)

Geoff

Geoff’s Stats Passions: A BJKS Podcast

When Benjamin Kuper-Smith kindly invited me to chat with him for his podcast I warned him he’d have trouble shutting me up. Maybe Ben felt that, but I felt we had a pretty interesting chat about lots of great (imho) stats issues. The podcast is here.

There’s an auto-generated transcript, and you can hover just below the moving sound line to see a control of audio replay speed–I find x1.25 or even x1.5 can be good.

Timestamps
0:00:00: A brief history of statistics, p-values, and confidence intervals <what, half an hour is ‘brief‘?>
0:32:02: Meta-analytic thinking
0:42:56: Why do p-values seem so random?
0:45:59: Are p-values and estimation complementary?
0:47:09: How do I know how many participants I need (without a power calculation)?
0:50:27: Problems of the estimation approach (big data)
1:00:08: A book or paper more people should read
1:02:50: Something Geoff wishes he’d learnt sooner
1:04:52: Advice for PhD students and postdocs

You can see my podcast is #82. Ben told me he was about to chat with Brian Nosek, I’m sure that will be a great podcast coming in a week or two. A few others I found especially interesting:

80. Simine Vazire: scientific editing, the purpose of journals, and the future of psychological science

55. Angelika Stefan: p-hacking, simulations, and Shiny Apps

54. Jessica Kay Flake: Schmeasurement, making stats engaging, and the Psychological Science Accelerator

53. Chris Chambers: Registered Reports, scheduled peer-review, and science without journals

13. Joe Hilgard: Scientific fraud, reporting errors, and effects that are too big to be true

Have a browse at BJKS. Happy listening,

Geoff

Simine Writes About Our Second Edition

A clear and accessible introduction to statistics, perfect for beginners. This book covers both the old and the new – giving students the fundamentals they need to understand their field, while equipping them with a more sophisticated understanding of the pros and cons of those established practices. The focus on open science and integration with statistical tools (e.g., JAMOVI) makes the book particularly useful for training future researchers.

That’s the endorsement of our forthcoming second edition by Simine Vazire. Bob and I are, of course, enormously appreciative of her generous words. I recently posted (here) about her appointment as the incoming Editor-in-Chief of Psychological Science–a fantastic development imho.

The bunch of balloons on the right sketch Simine’s main research interests. Who better to write about Open Science and what’s needed to help us all–and our students–do better science? Thanks Simine.

Geoff

Free Online APA Conference on Teaching Research Excellence in Psychology, December 14th 2023 9:00am to 2:30pm EST

If you want to wrap up your winter semester with an invigorating online conference on teaching research excellence, you’re in luck, as the APA’s Teaching Research Excellence conference will be held via Zoom on December 14th, 2023 from 9:00 am to 2:30pm EST.

You can register to attend for free: https://docs.google.com/forms/d/e/1FAIpQLSeJV0SWNBRi206hmjMe3s2vaGVF-y2ksgseFtwteOrxUnsP8Q/viewform

The full program is below. It includes talks from John Edlund (Research Director of Psi Chi and executive editor of JSP, associate editor at Collabra Psychology), Stephen Chew (director of the Culture & Family Development Lab at Wellesley College), and Bob. The conference is organized by Zane Zheng of Lasell University.