Teaching Statistics Using Web Applets–Including esci

Recognise the image? Not many handbooks sport a big colour picture from esci on its cover 🙂

Joe Rodgers gives us a super-useful new resource. Chapter 1 (follow the steps in the figure caption) is a wonderful reflection on lessons from his lifetime of teaching statistics, then a brief summary of every chapter in the book.

At the publisher’s site for this book, click ‘preview this book’ (button may take a few moment to appear) under the pic of the cover. You can scroll through contents, etc, to the full text of Chapter 1.

The cover image is from Chapter 9, a great contribution by Chip Reichardt, a long-time friend at the University of Denver. I first called there for a quick visit back in 2009, when I was early in the writing of what became my first book, UTNS, just after my first YouTube video of the dance of the p values, and several years before Open Science burst onto the scene.

Chip was an enthusiastic user of the original ESCI that I was developing in MS Excel. (The version of ESCI used in UTNS is still available here.)

A Treasure Trove of Web Applets for Statistics Teaching

…that’s what Chip’s great Chapter 9 describes, together with his sage advice for selecting and using applets for a wide range of teaching aims. He gives this link to a list of all the applet links, so any of those more than two dozen applets is only a couple of clicks away.

esci-web Animating the Dance of the r Values

Below is a slightly edited version of what Chip writes about the dance of the r values.

Whatever you do, don’t overlook the applet at:

https://esci.thenewstatistics.com/esci-dance-r.html

(Hat tip to Gordon Moore, creator of esci-web, referred to here and whence came the book’s cover image.)

Pro tip: click the ‘?‘ top right in the control panel, it turns green, and then see pop-out tool tips as you hover the mouse over a control.

A large population of scores is presented as faint grey dots (just visible in the cover image) in a scatterplot. The applet then draws a random sample of data from the population and shows these as dark blue dots in the scatterplot (with any outliers in red), so you can see where the sample of data falls within the larger population of scores. You can click to add the regression line and r value for that sample.

Click the Take Sample button to see a different random sample from the population. Click Run Stop to see a sequence of random sets of blue dots dancing–the light grey dots of the population not moving–with the speed of dancing set by the slider. See also the regression line and the sample r value dance.

Use the sliders in the upper yellow and green panels to set sample size N and population correlation ρ (rho). Explore the effect of changing the N and ρ values

I hesitate to single out one applet from among many other excellent applets, but this applet is simply magnificent in showing how scatterplots, regression lines, and correlations vary across random samples of data. (Thank you Chip!)

The Animated r Heap

A few clicks and you can watch the sequence of sample r values drop down the screen and form the r heap, the empirical sampling distribution of r.

Control panel at left. In the bottom panel the checkboxes turn on display of the lower right panel (the upper scatterplot shrinks) and turn on display of the green dots for sample r values, which accumulate to form the sampling distribution of the r values. Red dots are r values whose 95% CIs (not shown) would not capture the population correlation value (ρ = .30, set by the slider in the upper green panel at left) and marked in the figure by the vertical blue line. Pop-out tool tips are on (the ‘?’ top right in the control panel is green): note the mouse pointer and tool tip in the figure.

Chip’s Chapter 9 is not included in the preview, but it’s well worth finding in your library–lots of great advice for great dynamic teaching demos.

Thank you Chip for this terrific new teaching resource.

Thank you Joe for highlighting esci on the cover!

Geoff

Simine Vazire Drops In for Dinner!

Of course I had to wear my APS uniform. You can just see, over Simine’s shoulder, the stunning artwork, portions of which appear on the cover of all my books.

It was great to see Simine last night, and to hear about the latest from APS and the world of journal editing. Her great love is still journal editing–very fortunately for psychology and indeed all of science. But when she wants quiet time for writing she takes to the road, finding quiet corners to use her laptop in coffee shops or, I was surprised to hear, breweries!

This last week she has slowly driven down from Sydney, coffee shop by coffee shop, to spend time with Fiona Fidler and with the MetaMelb group at The University of Melbourne.

Plastic plate, after 30 years of
trips through the dishwasher

She dropped in for a tuna pasta dinner, which grandkids Lucy and Zoe always request. Simine ate off this StatPlay classic picture of the green heap, long since renamed the mean heap.

These days you can explore the dances and many other goodies in esci web.

Enjoy!

The mean heap, from esci web

Geoff

P.S. On a personal note, we recently had a busy month to mark my 80th birthday (yikes!)–wonderful chamber music, two big parties, and now, to cap it off, a visit from Simine!

The p Value Casino Is Open–For Significance Roulette!

Excel 2003, the best version ever, was enshittified by MicroSoft in the 2007 version, which was way slower and dropped many wonderful animation facilities :-(. Even vast efforts would not get my great Significance Roulette simulation running in the new version.

Two videos use my Excel 2003 version to explain: https://tiny.cc/SigRoulette1 and https://tiny.cc/SigRoulette2 or simply search at YouTube for ‘significance roulette’.

Now Bradley Dean has built Significance Roulette to run in your browser–as part of esci web.

Significance Roulette, after an initial study gives p = .01

It has taken me close to 20 years to build the following argument that leads to significance roulette, and which is summarised in the first half of our open access article Calin-Jageman & Cumming, 2024

  • For approaching a century numerous distinguished scholars, including philosophers of science, statisticians, and psychologists, have published cogent critiques of p values, significance testing, and how researchers across science use these.
  • Even so, a large proportion of researchers, teachers, journals, and granting bodies use null hypothesis significance testing (NHST)–often based on p < .05 or p < .01–as the standard for concluding whether or not an effect exists, whether or not a result is large or important. Despite this logic being wrong-headed in so many ways!
  • Devotion to NHST and p <.05 resembles an addiction–the researcher’s heroin. Rational argument is not sufficient to shake the addiction. Could a dramatic demonstration, perhaps persuading via the gut rather than the brain, shake this addiction?
  • A striking but little-known feature of p values is that they are highly unreliable–their sampling variability is astonishingly large. Replicate a study, exactly the same but with a new random sample, and expect to obtain replication p that can take just about any value between 0 and 1!
  • Jerry Lai in his lovely PhD studies took three converging approaches to find that a large proportion of published researchers in psychology, medicine, and statistics severely under-estimate the amount of variability in the p value with replication.
  • My first demonstration of p value variability was the dance of the p values, the first video of which dates from 2009. Search YouTube for ‘dance of the p values’ to find several videos. You can also play with the dances in esci web.
  • My second approach arose from study of the probability distribution of replication p, the p value obtained in a replication. My highly-cited 2008 article has details.
  • I define the p interval as the 80% prediction interval for replication p. It’s astonishingly long! For example, if an initial study obtains p = .05, the p interval is (.0002, .65), meaning an 80% chance of p within that interval and fully a 10% chance it falls below .0002 and 10% above .65. After p = .01 the interval is (.00001, .41). After initial p = .001 (*** highly statistically significant) there is fully a 1 in 6 chance a replication does not even achieve p < .05! After p = .20 ns there is a 1 in 3 chance a replication finds p < .05!
  • I take the probability distribution of replication p following initial p = .01 and divide the area under the curve, which represents probability, into 38 equal areas. I label each area with the p value in the centre of the interval. I have 38 p values, a few large and many small and very small, which accurately represent that probability distribution.
  • I scatter those 38 p values randomly around the 38 bins of a roulette wheel. Simply click to spin, wait for a few moments, and see the replication p you might have obtained from a replication. Much faster and cheaper than the hassle of raising a grant, recruiting participants, hassling with ethics approval… and collecting data!
  • At Significance Roulette in esci web you can click between initial p of .05 and .01. With initial p = .01, replication p values tend a little smaller, of course. But the striking thing is how widely spread the p values are! The distributions are pictured to left of the wheel. For example, switch from .05 to .01 and note a slightly smaller number of deep blue, deep trombone sound, despairing figures for p > .1 ns and slightly more bright red, triumphant trumpet blast, elated figures for p < .001 ***. Click SPIN, and note your quickened heart beat, sweaty palms, and that you are holding your breath–will you be despairing or elated–and in only a few seconds you’ll know!

Will this approach to tackling addiction via the gut be more effective than the decades of argument addressed at the cortex?

Best of luck at the p Value Casino!

Take-home messages

  • p values are unbelievably unreliable
  • Any p value could easily have been just about any other value
  • No p value deserves our trust 🙁
  • Simply don’t use p values, there are much better ways 🙂

Geoff

P.S. Enormous thanks to Bradley and Bob, who made it all happen.

Vale Michael Kubovy (1940-2025), Professor of Patterns

I don’t think I ever met Michael, but have long known his Gestalt perception work. I now discover he did so much more, especially as a pioneer in data analysis. He and I would have agreed on many, many things.

This post is courtesy Alex Holcombe, who wrote:

A tribute to Michael Kubovy

You can browse Michael’s books on perception here, and his dazzling ‘Psychology of Perspective and Renaissance Arthere.

During WW2 he fled with his parents from France to Portugal, then eventually Israel, where he completed his university education with cog psy royalty: his master’s with Daniel Kahneman, who initially hired him as a laboratory assistant after a chance encounter at a corner shop, and his doctorate with Amos Tversky, whom he met as his commander in the Israeli reserves. 

Among his notable inventions was an auditory analogue to the random-dot stereogram. This allowed a listener to hear a hidden melody with their “third ear” that was entirely undetectable by either ear alone.

Considering data analysis, Michael took delight in the pleasures of taking data seriously. This meant finding a way to visualise and explore data, a key interest of statistician John Tukey. He loved the early software Data Desk, which even allowed you to use sliders to interact with a data visualisation!

Michael lamented reliance on null hypothesis significance testing (NHST), a major cause of the replication crisis. He left, at his death, a sadly unfinished book designed to provide a solution: Use visualisation to evaluate quantitative models. Part of his description:

“We have tools to see if the residuals from the model are normally distributed. And so what we do is we teach the students to use certain tools in a routine way to see if data deviate in any important way from the model that they’re proposing. So we use graphics a lot and we show by example. And because this is a textbook based on R, and there will be R code all over the place, they will have examples of good practice. We’re not going to preach, but we’re going to give examples of what to do. And we hope that by osmosis and by practicing with our examples that we give in the text and on our website, we hope that people will acquire best practices in their data analysis work.”

Vale Michael Kubovy.

Geoff

PS It’s worth reading Alex’s full tribute 🙂

Estimation, Open Science, and Bob’s Wonderful New esci

Our open access article just released at https://doi.org/10.1002/ijop.13132:

Highlights

  • Three dramatisations of the enormous unreliability of the p value. Can these help weaken researchers’ addiction to NHST that has withstood more than half a century of cogent rational critiques?
  • Bob’s wonderful new open-source esci software with great estimation-based figures: See worked examples, and work along if you wish.

Abstract

We argue that researchers should test less, estimate more, and adopt Open Science practices. We outline some of the flaws of null hypothesis significance testing and take three approaches to demonstrating the unreliability of the p value. We explain some advantages of estimation and meta-analysis (“the new statistics”), especially as contributions to Open Science practices, which aim to increase the openness, integrity, and replicability of research. We then describe esci (estimation statistics with confidence intervals): a set of online simulations, and an R package for estimation that integrates into jamovi and JASP. This software provides (a) online activities to sharpen understanding of statistical concepts (e.g., “The Dance of the Means”); (b) effects sizes and confidence intervals for a range of study designs, largely by using techniques recently developed by Bonett; (c) publication-ready visualisations that make uncertainty salient; and (d) the option to conduct strong, fair hypothesis evaluation through specification of an interval null. Although developed specifically to support undergraduate learning through the 2nd edition of our textbook, esci should prove a valuable tool for graduate students and researchers interested in adopting the estimation approach. Further information is at https://thenewstatistics.com.

Figure 1. Significance roulette. If an initial study obtains p=.01, an exact replication–just the same but with an new sample–will obtain a p value drawn from the enormous spread of values on the wheel.

The Enormous Unreliability of p

This is the first time (1) the dance of the p values (search YouTube), (2) significance roulette (Figure 1; and search YouTube), and (3) p intervals (see the article) have all been described together in print. Significance roulette has been around for a while but this is its first outing in print. Alas, p values simply don’t deserve our trust. Enjoy the figures!

esci web

This component of esci is a set of simulations and tools by our colleague Gordon Moore that run in any browser. Explore the dances, play with sampling distributions, find critical values, and more.

esci for Data Analysis

Bob’s esci is an open-source package in R, which can be run in R, or within jamovi or (by December 2024) in JASP. We describe the wide range of measures and designs esci can analyse, including meta-analysis, and work through several examples. We emphasise figures that highlight uncertainty, especially by picturing confidence intervals.

Figure 2. Part of jamovi screen showing selection of the ‘Gender math IAT’ data file for opening.

We argue that p values, if used at all, are most valuable in the context of hypothesis evaluation based on an interval null hypothesis, and best understood with the help of an esci figure–see Figure 3.

Interactions can be challenging to understand and interpret; again esci provides figures designed to help–see Figure 4.

Pro Tip: Data Files Now in esci

The article advises download of jamovi-format data files (Gender math IAT.omvGender math IAT ma.omvCampus Involvement.omv and MeditationBrain.omv) from https://osf.io/uhwj2. Since the final version of the article was submitted Bob has integrated into esci all the data files used in ITNS2, including these four, so download from OSF is no longer needed.

Figure 3. esci figure for two independent groups. At left, the data points, means and 95% CIs for the two groups. The black triangle marks the difference between the means. This and its 90% and 95% CIs are shown on the difference axis at right. The two CIs allow test of the interval null hypothesis, the pink stripe.

Examples of Analyses by esci

Figures 3 and 4 are just two illustrations from the example esci analyses discussed in the article.

To open a data file within esci, click top left in jamovi, then click Open, Data Library, and scroll to see all the data files for ITNS2 arranged by chapter. Figure 2 shows selection of the first example file used in the article.

Figure 3 is an esci figure from a two independent groups analysis of the Gender math IAT file. The grey areas on the CIs are what we call plausibility curves. These illustrate variation in the plausibility, or relative likelihood, that values across and beyond the interval are the population value.

Figure 4 is one of the ways esci can display a 2 x 2 interaction–part of an RCT analysis of the MeditationBrain file.

If you wish, work along with the examples. The rich UI (user interface) of esci gives lots of scope to make figures look just as you want them–there’s advice about how to tweak your figures to look like those in the article.

As ever, we’d love to hear your comments on the new book and new software. Enjoy.

Figure 4. One way esci displays a 2 x 2 interaction. The difference in slope of the two lines indicates the size of the interaction. The fans of faint lines give a rough indication of the extent of uncertainty in estimating the slopes of the lines.

Geoff

Choosing a Textbook Cover Design

It’s a delicious moment when the publisher sends a number of options their graphic designer has dreamed up for the cover. Below are the options for the three books. In each case, can you pick our choice? Our choices are below—don’t scroll down yet… 

UTNS (2012) …at left.

ITNS1 (2017) …below.

ITNS2 (2024) …below.

The three sets, all framed expertly by Lindsay my wife, hang above my desk:

Our Choices …do you think we got it right?

UTNS (2012) …at left

Middle option

 ITNS1 (2017) …at right

Leftmost option

ITNS2 (2024) …below

Top right option, as below left. (We were offered just the other five but asked to see the bottom right design in the bottom left colours, so Routledge sent the top right option, which we chose.)

However, when we received our printed books, we discovered that Routledge had actually used a modification of our chosen design, as at right. Not exactly our choice, but not bad.

‘Treasure’: Claire’s Gorgeous Resin Artwork

Treasure, at left, by Claire Layman, 150 × 50cm, resin on stretched canvas. Claire is an internationally recognised artist, also a longtime friend.

Walk into our living room and be struck by the vibrancy and depth of colour of Treasure, so much more alive than any small printed copy can be.

Claire generously agreed that Treasure could be used on the cover of our three statistics textbooks. See the note on the copyright page of each.

At right, top to bottom:

UTNS (2012)

ITNS1 (2017)

ITNS2 (2024)

People often make comments like: “Looks like slices through some sort of stones”, “It’s wriggling things under a microscope!” or “Go snorkelling and see things like those?”

Claire’s response is “It’s an artwork, see it as you wish!” She mentions also that there’s no official top or bottom: hang it horizontally or vertically, either way up.

Besides looking great, is there any justification for it appearing on statistics books? People often make comments like: “There’s a pattern of those blues—oh, no there’s not”, or “Look, those ones sort-of alternate, but not quite”. Claire says that she often “starts to make a pattern, then breaks it”. That all sounds to me like trying to find some sort of regularity in randomness, which is one way of describing the aim of statistical inference: Can we identify a difference, or other pattern, lurking in the sampling variability, how large or strong is that pattern, and how confident can we be in our conclusion? That’s the central concern of our books.

Do you agree with the graphic designer’s choice of part-images from Treasure for the covers?

Geoff

NEXT: Choosing a cover design.     

Estimation in Neuroscience: Characterizing local circuits in the Inferior Colliculus

Here’s more evidence that estimation is catching on in neuroscience, a beautiful paper in from Silveira et al. (2023) in The Journal of Neuroscience. The paper characterizes local functional circuitry in the inferior colliculus (IC), a key processing station for auditory information ​(Silveira et al., 2023)​. The authors show that excitatory neurons in the IC can excite each as well as putative inhibitory neurons, meaning that there is both a local recurrent excitatory network as well as a possible feed-forward inhibitory loop. They show that they can produce recurrent bouts of excitation by stimulating excitatory neurons in the IC when GABA is blocked, and that neuropeptide Y can then limit/block these recurrent excitation. It’s lovely, careful work that helps map out the functional circuitry at an important auditory processing center. Best of all, the authors make extensive use of estimation, rolling their own Gardner-Altman plots with bootsrapped confidence intervals. Below, for example, is one of their figures showing that neuropeptide Y (NPY) limits the recurrent excitation produced during activation of a subclass of excitatory neurons in the IC. (Figure reproduced without permission :-().

I (Bob) am going to reach out to the authors to see why they ended up making their own estimation figures (they cite DABEST, but it looks like they didn’t end up using it), and if there are any ways esci could be updated to suit their needs.

  1. Silveira, M. A., Drotos, A. C., Pirrone, T. M., Versalle, T. S., Bock, A., & Roberts, M. T. (2023). Neuropeptide Y Signaling Regulates Recurrent Excitation in the Auditory Midbrain. Society for Neuroscience. doi: 10.1523/jneurosci.0900-23.2023

The Latest on esci: Geoff’s Talk to Turku

Yesterday I greatly enjoyed a zoom talk with the Open Science Community of Turku, Finland. Online were good folks from Finland, The Netherlands and I don’t know where else. Great questions at the end!

My slides and 3 data files (download the lot as a zip) are at osf.io/uhwj2 — you may need to log in to OSF but registering is simple.

The talk is up on YouTube here. I start at about 5.00 and esci at about 39.20.

I moved rapidly through dances, significance roulette, the plausibility curve on a confidence interval, esci web, Open Science, … to what’s new: the second edition of ITNS, due March ’24, AND the latest on esci.

Starting at Slide 61 there’s step-by-step detail of how to instal jamovi, instal esci into jamovi, and see the wide range of analyses esci offers, with examples. Plenty of guidance for anyone to get started using Bob’s great esci application. All open and free.

The Open Science Community of Turku boasts members from many disciplines and has the wonderful mission of “making Open Science the norm in research“. Way to go!

My thanks to Lydia Laninga-Wijnen, also Oskari Lahtinen, for the invitation and hosting.

Geoff

The Diamond Ratio (DR), Our Estimate of Heterogeneity: Now Published Online

I recently posted about our DR article being accepted by BJMSP. It has now been published online, here.

It’s behind a paywall, but here is the full pdf that we are allowed to share; it has only a few limitations, including no local file download.

Again, well done Max!

Enjoy,

Geoff