Teaching Statistics Using Web Applets–Including esci

Recognise the image? Not many handbooks sport a big colour picture from esci on its cover šŸ™‚

Joe Rodgers gives us a super-useful new resource. Chapter 1 (follow the steps in the figure caption) is a wonderful reflection on lessons from his lifetime of teaching statistics, then a brief summary of every chapter in the book.

At the publisher’s site for this book, click ‘preview this book’ (button may take a few moment to appear) under the pic of the cover. You can scroll through contents, etc, to the full text of Chapter 1.

The cover image is from Chapter 9, a great contribution by Chip Reichardt, a long-time friend at the University of Denver. I first called there for a quick visit back in 2009, when I was early in the writing of what became my first book, UTNS, just after my first YouTube video of the dance of the p values, and several years before Open Science burst onto the scene.

Chip was an enthusiastic user of the original ESCI that I was developing in MS Excel. (The version of ESCI used in UTNS is still available here.)

A Treasure Trove of Web Applets for Statistics Teaching

…that’s what Chip’s great Chapter 9 describes, together with his sage advice for selecting and using applets for a wide range of teaching aims. He gives this link to a list of all the applet links, so any of those more than two dozen applets is only a couple of clicks away.

esci-web Animating the Dance of the r Values

Below is a slightly edited version of what Chip writes about the dance of the r values.

Whatever you do, don’t overlook the applet at:

https://esci.thenewstatistics.com/esci-dance-r.html

(Hat tip to Gordon Moore, creator of esci-web, referred to here and whence came the book’s cover image.)

Pro tip: click the ‘?‘ top right in the control panel, it turns green, and then see pop-out tool tips as you hover the mouse over a control.

A large population of scores is presented as faint grey dots (just visible in the cover image) in a scatterplot. The applet then draws a random sample of data from the population and shows these as dark blue dots in the scatterplot (with any outliers in red), so you can see where the sample of data falls within the larger population of scores. You can click to add the regression line and r value for that sample.

Click the Take Sample button to see a different random sample from the population. Click Run Stop to see a sequence of random sets of blue dots dancing–the light grey dots of the population not moving–with the speed of dancing set by the slider. See also the regression line and the sample r value dance.

Use the sliders in the upper yellow and green panels to set sample size N and population correlation ρ (rho). Explore the effect of changing the N and ρ values

I hesitate to single out one applet from among many other excellent applets, but this applet is simply magnificent in showing how scatterplots, regression lines, and correlations vary across random samples of data. (Thank you Chip!)

The Animated r Heap

A few clicks and you can watch the sequence of sample r values drop down the screen and form the r heap, the empirical sampling distribution of r.

Control panel at left. In the bottom panel the checkboxes turn on display of the lower right panel (the upper scatterplot shrinks) and turn on display of the green dots for sample r values, which accumulate to form the sampling distribution of the r values. Red dots are r values whose 95% CIs (not shown) would not capture the population correlation value (ρ = .30, set by the slider in the upper green panel at left) and marked in the figure by the vertical blue line. Pop-out tool tips are on (the ‘?’ top right in the control panel is green): note the mouse pointer and tool tip in the figure.

Chip’s Chapter 9 is not included in the preview, but it’s well worth finding in your library–lots of great advice for great dynamic teaching demos.

Thank you Chip for this terrific new teaching resource.

Thank you Joe for highlighting esci on the cover!

Geoff

Vale Michael Kubovy (1940-2025), Professor of Patterns

I don’t think I ever met Michael, but have long known his Gestalt perception work. I now discover he did so much more, especially as a pioneer in data analysis. He and I would have agreed on many, many things.

This post is courtesy Alex Holcombe, who wrote:

A tribute to Michael Kubovy

You can browse Michael’s books on perception here, and his dazzling ‘Psychology of Perspective and Renaissance Arthere.

During WW2 he fled with his parents from France to Portugal, then eventually Israel, where he completed his university education with cog psy royalty: his master’s with Daniel Kahneman, who initially hired him as a laboratory assistant after a chance encounter at a corner shop, and his doctorate with Amos Tversky, whom he met as his commander in the Israeli reserves.Ā 

Among his notable inventions was an auditory analogue to the random-dot stereogram. This allowed a listener to hear a hidden melody with their “third ear” that was entirely undetectable by either ear alone.

Considering data analysis, Michael took delight in the pleasures of taking data seriously. This meant finding a way to visualise and explore data, a key interest of statistician John Tukey. He loved the early software Data Desk, which even allowed you to use sliders to interact with a data visualisation!

Michael lamented reliance on null hypothesis significance testing (NHST), a major cause of the replication crisis. He left, at his death, a sadly unfinished book designed to provide a solution: Use visualisation to evaluate quantitative models. Part of his description:

“We have tools to see if the residuals from the model are normally distributed. And so what we do is we teach the students to use certain tools in a routine way to see if data deviate in any important way from the model that they’re proposing. So we use graphics a lot and we show by example. And because this is a textbook based on R, and there will be R code all over the place, they will have examples of good practice. We’re not going to preach, but we’re going to give examples of what to do. And we hope that by osmosis and by practicing with our examples that we give in the text and on our website, we hope that people will acquire best practices in their data analysis work.”

Vale Michael Kubovy.

Geoff

PS It’s worth reading Alex’s full tribute šŸ™‚

ITNS2 Data Sets Are Now Available Within esci

All the data sets used in examples and exercises in the new second edition of ITNS are now available within esci in jamovi, esci in R, and soon within esci in JASP.

Here’s how to open one of those data sets in jamovi.

  1. Open jamovi with the esci module installed–meaning you can see the esci icon in the top bar, as pictured.
  2. Click three lines, top left. Panel opens.
  3. Click Open, Data Library.
  4. See 47 data files, ordered by chapter of first mention in the book.
  5. Click on your choice of data set.

In the book, ignore instructions to go to the book’s Companion Website to download a data set and save locally before opening in jamovi.

Thanks Bob, life is now easier!

Geoff

Estimation, Open Science, and Bob’s Wonderful New esci

Our open access article just released at https://doi.org/10.1002/ijop.13132:

Highlights

  • Three dramatisations of the enormous unreliability of the p value. Can these help weaken researchers’ addiction to NHST that has withstood more than half a century of cogent rational critiques?
  • Bob’s wonderful new open-source esci software with great estimation-based figures: See worked examples, and work along if you wish.

Abstract

We argue that researchers should test less, estimate more, and adopt Open Science practices. We outline some of the flaws of null hypothesis significance testing and take three approaches to demonstrating the unreliability of the p value. We explain some advantages of estimation and meta-analysis (ā€œthe new statisticsā€), especially as contributions to Open Science practices, which aim to increase the openness, integrity, and replicability of research. We then describe esci (estimation statistics with confidence intervals): a set of online simulations, and an R package for estimation that integrates into jamovi and JASP. This software provides (a) online activities to sharpen understanding of statistical concepts (e.g., ā€œThe Dance of the Meansā€); (b) effects sizes and confidence intervals for a range of study designs, largely by using techniques recently developed by Bonett; (c) publication-ready visualisations that make uncertainty salient; and (d) the option to conduct strong, fair hypothesis evaluation through specification of an interval null. Although developed specifically to support undergraduate learning through the 2nd edition of our textbook, esci should prove a valuable tool for graduate students and researchers interested in adopting the estimation approach. Further information is at https://thenewstatistics.com.

Figure 1. Significance roulette. If an initial study obtains p=.01, an exact replication–just the same but with an new sample–will obtain a p value drawn from the enormous spread of values on the wheel.

The Enormous Unreliability of p

This is the first time (1) the dance of the p values (search YouTube), (2) significance roulette (Figure 1; and search YouTube), and (3) p intervals (see the article) have all been described together in print. Significance roulette has been around for a while but this is its first outing in print. Alas, p values simply don’t deserve our trust. Enjoy the figures!

esci web

This component of esci is a set of simulations and tools by our colleague Gordon Moore that run in any browser. Explore the dances, play with sampling distributions, find critical values, and more.

esci for Data Analysis

Bob’s esci is an open-source package in R, which can be run in R, or within jamovi or (by December 2024) in JASP. We describe the wide range of measures and designs esci can analyse, including meta-analysis, and work through several examples. We emphasise figures that highlight uncertainty, especially by picturing confidence intervals.

Figure 2. Part of jamovi screen showing selection of the ‘Gender math IAT’ data file for opening.

We argue that p values, if used at all, are most valuable in the context of hypothesis evaluation based on an interval null hypothesis, and best understood with the help of an esci figure–see Figure 3.

Interactions can be challenging to understand and interpret; again esci provides figures designed to help–see Figure 4.

Pro Tip: Data Files Now in esci

The article advises download of jamovi-format data files (Gender math IAT.omv,Ā Gender math IAT ma.omv,Ā Campus Involvement.omvĀ andĀ MeditationBrain.omv) from https://osf.io/uhwj2. Since the final version of the article was submitted Bob has integrated into esci all the data files used in ITNS2, including these four, so download from OSF is no longer needed.

Figure 3. esci figure for two independent groups. At left, the data points, means and 95% CIs for the two groups. The black triangle marks the difference between the means. This and its 90% and 95% CIs are shown on the difference axis at right. The two CIs allow test of the interval null hypothesis, the pink stripe.

Examples of Analyses by esci

Figures 3 and 4 are just two illustrations from the example esci analyses discussed in the article.

To open a data file within esci, click top left in jamovi, then click Open, Data Library, and scroll to see all the data files for ITNS2 arranged by chapter. Figure 2 shows selection of the first example file used in the article.

Figure 3 is an esci figure from a two independent groups analysis of the Gender math IAT file. The grey areas on the CIs are what we call plausibility curves. These illustrate variation in the plausibility, or relative likelihood, that values across and beyond the interval are the population value.

Figure 4 is one of the ways esci can display a 2 x 2 interaction–part of an RCT analysis of the MeditationBrain file.

If you wish, work along with the examples. The rich UI (user interface) of esci gives lots of scope to make figures look just as you want them–there’s advice about how to tweak your figures to look like those in the article.

As ever, we’d love to hear your comments on the new book and new software. Enjoy.

Figure 4. One way esci displays a 2 x 2 interaction. The difference in slope of the two lines indicates the size of the interaction. The fans of faint lines give a rough indication of the extent of uncertainty in estimating the slopes of the lines.

Geoff

Fun with esci in R: The simple two-group design

esci is now available as a module in jamovi and as a package in R (JASP coming soon, hopefully). Let’s have some fun with esci in R!

We’ll start with a simple two-group design. Specifically, we’ll use data from Experiment 4 of ​(Kardas & O’Brien, 2018)​. In this study, participants watched a video explaining how to do a simple mirror-tracing task ​(Cusack, Vezenkova, Gottschalk, & Calin-Jageman, 2015)​. Participants were randomly assigned to watch the training video either 1 time or 20 times. They then predicted how they would perform on the task (0-100%) and then completed the task (0-100%). Kardas & O’Brien found that watching the training video repeatedly boosted confidence (predicted scores) but not performance.

Opening the data – R

If you haven’t installed esci yet, you can do so with:

install.packages("esci")

Once installed, we will load esci into memory and then we will store the Kardas & O’Brien data set bundled in esci, giving it the name mydata


library(esci)
mydata <- esci::data_kardas_expt_4

Analyze the data in R

We are going to analyze the effect of video Exposure on Prediction scores. We can do this with the estimate_mdiff_two command. We’ll want to store the result, so tell R to store it in a new variable called estimate.


estimate <- esci::estimate_mdiff_two(
  data = mydata,
  outcome_variable = Prediction,
  grouping_variable = Exposure,
  conf_level = 0.95,
  assume_equal_variance = TRUE
)

Note that we’ve decided to assume equal variance… but it’s probably a better default not to do this.. and it’s easy enough to change the command by setting assume_equal_variance to FALSE.

Inspect the result

What we get back in R is a list, a complex object that contains other objects. You can inspect this object in lots of different ways, but lets try listing the objects it contains:

names(estimate)
 [1] "properties"                      "es_mean_difference_properties"  
 [3] "es_mean_difference"              "es_median_difference"           
 [5] "es_median_difference_properties" "es_smd_properties"              
 [7] "es_smd"                          "es_mean_ratio"                  
 [9] "es_mean_ratio_properties"        "es_median_ratio"                
[11] "es_median_ratio_properties"      "overview"                       
[13] "raw_data"             

We can see that our results has properties, and then a bunch of different objects that start with es — that’s short for effect size. We get a mean difference, a median difference, an smd, a mean ratio, and a median ratio. Many of these have their own properties as well. Finally, we also get an overview and raw_data.

Let’s see the overview:

> estimate$overview
  outcome_variable_name grouping_variable_name grouping_variable_level     mean  mean_LL  mean_UL median
1            Prediction               Exposure                       1 56.37795 52.90820 59.84770     60
2            Prediction               Exposure                      20 67.76224 64.49236 71.03212     71
  median_LL median_UL       sd min max q1   q3   n missing  df  mean_SE median_SE
1  53.57284  66.42716 22.07273   0 100 40 71.5 127       0 268 1.762318  3.279225
2  66.12580  75.87420 17.66669   0 100 59 81.0 143       0 268 1.660803  2.486881

You can see that overview is a table — it lists each group found in the data (1x and 20x exposure) and provides basic descriptive statistics: mean with confidence interval, median with confidence interval, standard deviation, etc.

Let’s take a look at the es_mean_difference table:

> estimate$es_mean_difference
        type outcome_variable_name grouping_variable_name effect effect_size        LL
1 Comparison            Prediction               Exposure     20    67.76224 64.492358
2  Reference            Prediction               Exposure      1    56.37795 52.908204
3 Difference            Prediction               Exposure 20 ‒ 1    11.38429  6.616553
        UL       SE  df     ta_LL    ta_UL
1 71.03212 1.660803 268 65.020985 70.50349
2 59.84770 1.762318 268 53.469143 59.28676
3 16.15202 2.421576 268  7.387331 15.38124

You can see that this table gives us the mean and confidence interval of the 20x group, of the 1x group, and of the difference between them, reporting (in row 3) the contrast between the 20x and 1x group. The 1x group, in this case is the reference group — we express the effect size relative to the 1x group. The mean difference in prediction scores is 11.38 95% CI [6.6, 16.15]. We also get the standard error, degrees of freedom, and the confidence interval at two alpha (90% CI in this case). Clearly, watching the instructional video make a pretty big difference in predictions, it boosted confidence by over 10 points on a 0-100 scale in this sample! There is some uncertainty about the size of the effect, but overall, it seems clear that more video exposure leads to more confidence.

Notice that we have some other ways of expressing the effect size. For one, we can examine median differences–probably a better idea in most cases in psychology, but not widely done.

> estimate$es_median_difference
        type outcome_variable_name grouping_variable_name effect effect_size        LL       UL       SE
1 Comparison            Prediction               Exposure     20          71 66.125803 75.87420 2.486881
2  Reference            Prediction               Exposure      1          60 53.572837 66.42716 3.279225
3 Difference            Prediction               Exposure 20 ‒ 1          11  2.933636 19.06636 4.115567
      ta_LL    ta_UL
1 66.909445 75.09055
2 54.606155 65.39385
3  4.230494 17.76951

This table is setup similarly to the es_mean_difference — we again get each group and the contrast between them. There is more uncertainty here (a difference of 11 points with a 95% CI [2.9, 19.06])… the data are consistent with a large median difference but also with a fairly small median difference of just 2.9 points (and valuers near the CI boundary are not very different in their compatibility with the data). So we’d still want to be cautious about concluding there is a meaningful median difference.

Want more ways to express this? Of course! We can also thing about the ratios between the group means or medians. Here’s the ratio of medians:

> estimate$es_median_ratio
  outcome_variable_name grouping_variable_name effect effect_size       LL       UL comparison_median
1            Prediction               Exposure 20 / 1    1.183333 1.037629 1.349498                71
  reference_median
1               60

The 20x group had a median 1.18x the 1x group, but the CI is broad [1.037, 1.349].. so somewhere between a very small to very large increase in median is compatible with this data.

And, of course, psychologists remain a bit obsessed with Cohen’s d. So let’s look at the es_smd table:

> estimate$es_smd
  outcome_variable_name grouping_variable_name effect effect_size       LL        UL numerator denominator
1            Prediction               Exposure 20 ‒ 1   0.5716119 0.327274 0.8149238  11.38429    19.86031
SE df d_biased 1 0.1244027 268 0.5732178

This is a fairly large effect: d = 0.57 95% CI [0.33, 0.81] and the confidence interval is fairly narrow — we could fairly easily plan a sensitive follow-up study to help confirm and better characterize this effect.

But wait, there are lots of approaches to Cohen’s d… what is the denominator that was used and what flavor of Cohen’s d have we produced? Take a look at es_smd_properties to find out.

> estimate$es_smd_properties
$message This standardized mean difference is called d_s because the standardizer used was s_p. d_s has been corrected for bias. Correction for bias can be important when df < 50. See the rightmost column for the biased value.

Ah, so this is ds – because it used the pooled standard deviation. If we had set assume_equal_variance to FALSE we’d have obtained d_avg, which uses the average of the group standard deviations as the normalizer.

Visualizations

We don’t just want a bunch of tables… let’s see this data.

This is easy in esci, we just pass our stored result (estimate) to an appropriate plot function. In this case, we’ll use plot_mdiff to visualize a mean or median difference:

esci::plot_mdiff(
  estimate,
  effect_size = "mean"
)

and we get this beautiful figure:

Which we can then customize to our heart’s content (it’s a ggplot2 plot object).

Want to see the median difference instead? Here we go:

esci::plot_mdiff(
  estimate,
  effect_size = "median"
)

and we get:

Evaluating a Hypothesis

Although Kardas & O’Brien conducted several studies on video exposure, this was the first study they conducted using mirror tracing as the performance task. Therefore, they probably didn’t yet have a clear quantitative prediction to test — they weren’t really ready for hypothesis testing. Imagine, though, that you are going to conduct a replication study. Based on Kardas & O’Brien, you believe increased video exposure produces a substantive change in confidence, and you decide to define this as at least a 5 point difference in means. In other words, you’re specifying an interval null. The skeptic’s hypothesis (the null hypothesis) is that any difference in confidence will be negligible (< 5 point difference) and your hypothesis is that it will be substantive (>5 point difference).

We can visualize your prediction against the results by tweaking our call to plot_mdiff just a bit:

esci::plot_mdiff(
  estimate,
  effect_size = "mean", 
  rope = c(-5, 5)
)

We’ve passed a two-element vector that defines the interval null. This is called a ROPE or region of practical equivalence. We’ve defined the rope using the concatenate function in R which creates vectors: c(-5, 5) — that means create a vector with elements -5 and 5 and send that to the function where it is expecting a ROPE to be defined.

Here’s what we get:

You can see the ROPE shaded in in red and pink, and you can visually compare the results with the predictions of you and the skeptic. The rules for declaring victory or simple: if the whole CI of the result is inside the ROPE, the skeptic wins, if the whole CI is outside, you win, and if there is overlap there is a draw. In this case, we can see that the CI on the difference is fully outside the ROPE (though not by a ton). If the ROPE had really been established a priori and a sensitive experiment designed to test the predictions, we’d now have a strong confirmatory indication that their is, indeed, a substantive effect of video exposure on confidence (well, strong statistical evidence… we’d still need to think about the internal and external validity of our study and the extent to which this supports our claim as well).

Want to conduct the hypothesis test a bit more formally? esci can help with the test_mdiff function, which takes arguments very similar to what we passed to plot_mdiff:

esci::test_mdiff(
  estimate,
  effect_size = "mean",
  rope = c(-5, 5)
)

We get back a complex object, but one of its components is a table called interval_null which has this content:

$interval_null
                    test_type outcome_variable_name effect          rope
1 Practical significance test            Prediction 20 ‒ 1 (-5.00, 5.00)
  confidence                                                       CI
1         95 95% CI [6.616553, 16.15202]\n90% CI [7.387331, 15.38124]
              rope_compare p_result
1 95% CI fully outside H_0 p < 0.05 conclusion significant 1 At α = 0.05, conclude μ_diff is substantive TRUE

Viola!

In this example, we conducted a hypothesis test on mean differences, but we could just as easily work with median differences just by changing the effect_size argument to “median”. How cool is it to be able to conduct interval null tests of differences in median!? Think how sophisticated you will feel!

Conclusions

We’ve taken a quick tour of analyzing a two group design in esci.

esci is still in development. I expect the visualization functions, like plot_mdiff, to still change a bit. But the overall workflow should hopefully be stable and sensible: you generate an estimate with an estimate_ function, you can then visualize it (plot_ functions) and/or evaluate a hypothesis with it (test_ functions). The estimate_function produces complex lists with all the results you need: overview table, various es_ tables reporting different effect sizes, and _properties lists with all the nitty-gritty details. And that’s that!

  1. Cusack, M., Vezenkova, N., Gottschalk, C., & Calin-Jageman, R. J. (2015). Direct and Conceptual Replications of Burgmer &amp; Englich (2012): Power May Have Little to No Effect on Motor Performance (J. M. Haddad, Ed.). Public Library of Science (PLoS). doi: 10.1371/journal.pone.0140806
  2. Kardas, M., & O’Brien, E. (2018). Easier Seen Than Done: Merely Watching Others Perform Can Foster an Illusion of Skill Acquisition. SAGE Publications. doi: 10.1177/0956797617740646

The Latest on esci: Geoff’s Talk to Turku

Yesterday I greatly enjoyed a zoom talk with the Open Science Community of Turku, Finland. Online were good folks from Finland, The Netherlands and I don’t know where else. Great questions at the end!

My slides and 3 data files (download the lot as a zip) are at osf.io/uhwj2 — you may need to log in to OSF but registering is simple.

The talk is up on YouTube here. I start at about 5.00 and esci at about 39.20.

I moved rapidly through dances, significance roulette, the plausibility curve on a confidence interval, esci web, Open Science, … to what’s new: the second edition of ITNS, due March ’24, AND the latest on esci.

Starting at Slide 61 there’s step-by-step detail of how to instal jamovi, instal esci into jamovi, and see the wide range of analyses esci offers, with examples. Plenty of guidance for anyone to get started using Bob’s great esci application. All open and free.

The Open Science Community of Turku boasts members from many disciplines and has the wonderful mission of “making Open Science the norm in research“. Way to go!

My thanks to Lydia Laninga-Wijnen, also Oskari Lahtinen, for the invitation and hosting.

Geoff

2nd Edition Now With Routledge!

Bob and I are delighted to report that we’ve submitted ITNS2 to the publisher. Routledge say to expect it ‘early in 2024’. We’re hoping they may have something in time for faculty making textbook decisions for the N. Hemisphere Spring Semester.

Not only a new cover design, but also: Bob’s wonderful esci 1.0.1 is now available in jamovi, with a new name and new logo.

esci is now Estimation Statistics with Confidence Intervals, and the logo is:

The floating difference axis is a worthy logo because it’s so central to our estimation approach, and appears in so many esci analyses.

Bob will have more to say about esci 1.0.1 but here’s how to start:

  • At the jamovi downloads page, click the latest version for your choice of Windows, macOS, or Linux. (Currently 2.4.5 for Windows.) Install jamovi.
  • In jamovi, click the large plus sign, top right, choose ‘Manage Installed’, scroll through the ‘Available’ modules to find esci 1.0.1, click to install.
  • Enjoy seeing a new menu appear: esci with its logo.

We hope all these goodies serve you and your students well and we look forward to hearing about your experiences.

Geoff

The Diamond Ratio (DR), Our Estimate of Heterogeneity: Now Published Online

I recently posted about our DR article being accepted by BJMSP. It has now been published online, here.

It’s behind a paywall, but here is the full pdf that we are allowed to share; it has only a few limitations, including no local file download.

Again, well done Max!

Enjoy,

Geoff

The Diamond Ratio (DR), Our Estimate of Heterogeneity: Accepted for Publication

I’m excited to report that Max Cairns’s PhD work on the Diamond Ratio (DR) has been accepted for publication by the British Journal of Mathematical and Statistical Psychology. The preprint of the final accepted version is here.

Our preprint, as accepted by BJMSP

The Original DR Blog Post

…was in 2018: Measuring Heterogeneity in Meta-Analysis: The Diamond Ratio (DR)

It includes a brief intro to the Fixed Effect (FE) and Random Effects (RE) models for meta-analysis, and to heterogeneity. When there’s heterogeneity, the diamond depicting the 95% CI on the result of an RE meta-analysis is likely to be longer than the FE diamond.

The DR is simply the ratio of those two diamond lengths. If DR = 1 there’s little or no heterogeneity; DR of, say, 1.5 suggests moderate heterogeneity. Larger DR, more heterogeneity.

Bob and I introduced the DR in Chapter 9 of ITNS. We wanted to discuss heterogeneity, while avoiding the complexities of conventional estimates of heterogeneity (Q, I2, Τ2). If both the RE and FE diamonds are displayed at the bottom of a forest plot, the DR can easily be eyeballed (see figure below).

A CI on the DR

Last year I posted about Max’s success in developing a very good approximate CI on the DR:

A Confidence Interval for the Diamond Ratio: Estimation of Heterogeneity in Meta-Analysis

The Paper Accepted for Publication

It’s easy to eyeball the DR—the ratio of the lengths of the two red diamonds in this figure from the preprint:

Forest plot summarising a meta-analysis performed on data in Figure 9.2 of ITNS. Both RE and FE diamonds are displayed, in red. The DR and its CI are reported (lower left) to be 1.40 [1.00, 3.09]. Eyeballing the ratio of the lengths of the diamonds agrees with that reported value of DR.

The preprint describes Max’s investigations of seven (!) approaches to calculating a CI for DR. These are all approximations, but, as the preprint reports, Max carried out extensive simulations to evaluate all of these. His main focus was on coverage. He found that the ā€˜Sub-Q’ approach has excellent coverage (very close to 95%) across a very wide range of situations.

Max’s short description of this CI is that it is the ā€œSubstitution CI with the Q-profile Ļ„2 interval estimatorā€. No, that’s not totally clear to me either, but see the preprint for the full story, including the impressive (imho) range of simulation results. There are links to all the data and results, and to R software for calculating the CI on DR.

The red line below the RE diamond in the figure represents the length of the 95% prediction interval (PI) for the population effect sizes estimated by different individual studies. In other words, it indicates the likely extent of spread of these true effect sizes. It’s a further way to visualise the likely extent of heterogeneity. The length of the PI is 0.29, as reported in the figure.

esci in jamovi

Bob has included calculation of the CI on DR in the current beta of esci in jamovi, available here.

Visualisation is Understanding

Well, very often it’s a big help, and not only for beginning students. Visualisation is a focus all through esci and ITNS. We hope DR visualisation will help students and even researchers achieve a better understanding of heterogeneity.

As usual, it’s highly valuable to have a CI as well as a point estimate—in this case, of heterogeneity. Unless k, the number of studies in the meta-analysis, is quite large, the CI on DR is likely to be long, as it is in the figure above. That CI extends from the minimum, 1, to more than 3. With only k = 10 studies, all we can say is that heterogeneity is most likely between zero and very large. In other words, the amount of heterogeneity could be pretty much anything. That’s unfortunate, but it’s better to know this than be misled by any seemingly precise point estimate.

Well done Max! (And supervisor, Luke.)

Geoff

Teaching Statistics: Great Talks and Resources Now Online

This conference (as in pic below) was planned as a gathering of, maybe, 30 folks. But then it had to go online and the organisers (Kevin Peters and Fegal O’Hagan, of Trent University, and Rob Cribbie, of York University) found themselves scrambling to expand their Zoom limit, as more than 600 folks from around the world registered. Success!

You can see here the list of speakers, the abstracts, and video recordings of most of the talks, and links to slides and other resources provided by many of the speakers.

Our talk

Teaching the New Statistics, Now With Better Software

Bob’s and my talk was the last—at traditional conferences the grave-yard slot when many have already fled to the airport. It was at 6am in the winter dark for me, but that was OK. At the site, scroll to the end to see our links, then click the down arrow to see our abstract.

My job was to demonstrate the great new software: esci web by Gordon Moore, and esci in jamovi by Bob. Both are freely accessible from the esci menu at our site. For esci web, search at our site for ā€˜Gordon’ to find four blog posts.

Some starting points in the video of our talk:

    10.47 esci web

    25.35 dance of the p values, in esci web

    35.30 variability of p values with replication: Significance Roulette

    40.00 esci in jamovi

        41.40 two independent groups example

        45.50 single group example

        47.41 meta-analysis

        50.02 the diamond ratio (heterogeneity in a meta-analysis)

        50.07 interaction in a two-way design

The conference day ran from my 11 pm to 7 am, so I managed to catch only a few other talks. These included:

Andy Field

Teaching statistics: Damnation and deliverance

Lively and engaging, this talk included a dazzling array of videos in various punk and gothic styles (I may have those terms a bit wrong…) designed to highlight statistical ideas. Definitely different, even if, I suspect, not to everyone’s taste. However, given the wild success of his statistics textbooks, Andy must be doing many things right.

Christopher Green

The Replication Crisis: What should we teach to undergraduates, and when?

Interesting discussion of Open Science—why it’s needed and what we should teach about it. Starting at 21.55 are some arguments for not teaching undergraduates about meta-analysis. I’m not persuaded, while fully agreeing that selective publication (thank you NHST) is an enormous problem for meta-analysis—and science.

E.J. (Eric-Jan Wagenmakers)

Tips and tricks for teaching Bayesian statistics

E.J. was in good form, as lively and compelling as ever. I took my first course on Bayesian statistics in 1965 and have read books and attended many workshops since. The logic, and the match with how human cognition works have always appealed. I’ve kept looking out for simple and practical ways to introduce Bayesian methods in the intro statistics course for psychology students. Ideally, these should also help seasoned researchers brought up in the NHST tradition (ugh!) to understand and adopt Bayesian methods. I’m still looking, which is why I’ve focused on traditional frequentist CIs as the practical way forward, at least at first. I’d be more than happy to see Bayesian estimation, based on credible intervals, and Bayesian modelling much more widely used.

At 5.30 see E.J.’s book—a free download—Bayesian thinking for toddlers, which presents an ingenious and extremely simple way to introduce the core Bayesian idea of using evidence to update belief.

From 20.42 E.J. is talking about JASP, the wonderful SPSS-killing open-source statistics application that his team has been developing for some time. It supports frequentist as well as Bayesian methods. Bob is working with the team towards have esci available within JASP before too long.

At 45.40 E.J. recommends Bayesian Statistics the Fun Way by Will Kurt. I’ve sent off for it—perhaps this will give me the easy way in that I’ve long been seeking?

Themed Session: The Flipped Classroom

Bob and I have reports that some instructors are successfully using ITNS and its online resources for a flipped-classroom approach. That’s been great to hear—flipping may be becoming widespread, especially after 2020, and we hope that ITNS2 and its materials will be even more flip-friendly. So I was especially interested to see these three talks from flipsters at McMaster and York Universities.

Flipping inferential statistics

The flipped classroom improves performance in introductory statistics: Early evidence from a systematic review and meta-analysis

This is the one talk of the three for which a video is available, at least at present. The meta-analysis included 11 studies comparing flipped with lecture formats. The overall mean effect size advantage for flipped was g = 0.40, with part of that attributable to use of regular quizzes.

No tutorials, no problem: The inspiration, planning, and execution of flipping a Statistics II course

Wesley Burr

Teaching Reproducibility and Replicability in Statistics

This talk was just before mine; I caught the last part. Great enthusiasm and engagement. Lots on Open Science. What seemed like excellent advice on using R from the start, via R Markdown. Pointers to resources.

Finally

I highly recommend browsing the abstracts and at least dipping in to the videos. Please make a comment below on any you find especially interesting or useful. Thanks.

Geoff

P.S. Congratulations and thanks to Kevin, Fegal, and Rob for their initiative and vast amount of work. There are already discussions about when another such conference might be organized.