Teaching Statistics Using Web Applets–Including esci

Recognise the image? Not many handbooks sport a big colour picture from esci on its cover 🙂

Joe Rodgers gives us a super-useful new resource. Chapter 1 (follow the steps in the figure caption) is a wonderful reflection on lessons from his lifetime of teaching statistics, then a brief summary of every chapter in the book.

At the publisher’s site for this book, click ‘preview this book’ (button may take a few moment to appear) under the pic of the cover. You can scroll through contents, etc, to the full text of Chapter 1.

The cover image is from Chapter 9, a great contribution by Chip Reichardt, a long-time friend at the University of Denver. I first called there for a quick visit back in 2009, when I was early in the writing of what became my first book, UTNS, just after my first YouTube video of the dance of the p values, and several years before Open Science burst onto the scene.

Chip was an enthusiastic user of the original ESCI that I was developing in MS Excel. (The version of ESCI used in UTNS is still available here.)

A Treasure Trove of Web Applets for Statistics Teaching

…that’s what Chip’s great Chapter 9 describes, together with his sage advice for selecting and using applets for a wide range of teaching aims. He gives this link to a list of all the applet links, so any of those more than two dozen applets is only a couple of clicks away.

esci-web Animating the Dance of the r Values

Below is a slightly edited version of what Chip writes about the dance of the r values.

Whatever you do, don’t overlook the applet at:

https://esci.thenewstatistics.com/esci-dance-r.html

(Hat tip to Gordon Moore, creator of esci-web, referred to here and whence came the book’s cover image.)

Pro tip: click the ‘?‘ top right in the control panel, it turns green, and then see pop-out tool tips as you hover the mouse over a control.

A large population of scores is presented as faint grey dots (just visible in the cover image) in a scatterplot. The applet then draws a random sample of data from the population and shows these as dark blue dots in the scatterplot (with any outliers in red), so you can see where the sample of data falls within the larger population of scores. You can click to add the regression line and r value for that sample.

Click the Take Sample button to see a different random sample from the population. Click Run Stop to see a sequence of random sets of blue dots dancing–the light grey dots of the population not moving–with the speed of dancing set by the slider. See also the regression line and the sample r value dance.

Use the sliders in the upper yellow and green panels to set sample size N and population correlation ρ (rho). Explore the effect of changing the N and ρ values

I hesitate to single out one applet from among many other excellent applets, but this applet is simply magnificent in showing how scatterplots, regression lines, and correlations vary across random samples of data. (Thank you Chip!)

The Animated r Heap

A few clicks and you can watch the sequence of sample r values drop down the screen and form the r heap, the empirical sampling distribution of r.

Control panel at left. In the bottom panel the checkboxes turn on display of the lower right panel (the upper scatterplot shrinks) and turn on display of the green dots for sample r values, which accumulate to form the sampling distribution of the r values. Red dots are r values whose 95% CIs (not shown) would not capture the population correlation value (ρ = .30, set by the slider in the upper green panel at left) and marked in the figure by the vertical blue line. Pop-out tool tips are on (the ‘?’ top right in the control panel is green): note the mouse pointer and tool tip in the figure.

Chip’s Chapter 9 is not included in the preview, but it’s well worth finding in your library–lots of great advice for great dynamic teaching demos.

Thank you Chip for this terrific new teaching resource.

Thank you Joe for highlighting esci on the cover!

Geoff

Psychologists Help Society in so Many Ways

Portrait of Ron Cumming, aeronautical engineer, psychologist, pioneer (with Ross Day) of human factors in Australia

Ten Years Applying Psychological Science Inside the U.K. Government is fascinating. It prompted this post and is at 6. below, but first a back story.

TL;DR. Main points:

  1. Ron Cumming, my father, pioneered human factors in Australia in the 1960s.
  2. Human factors, or user-centred design (UCD), is the design of devices and systems to be safe, easy, and effective for users.
  3. Ron and colleagues persuaded Victorian politicians to pass in 1970 the world’s first compulsory seat belt laws–road fatalities dropped immediately.
  4. Thaler & Sunstein’s Nudge gave examples of how simple changes in messages and systems can help people make better choices.
  5. I taught UCD at La Trobe for many years. Students had great fun identifying easy ways to improve the design of everyday things simply by watching and talking to real users. Some tell me they were inspired to go on and work in the field.
  6. In her fascinating column Ten Years Applying Psychological Science Inside the U.K. Government Carla Groom describes how nudge and UCD ideas can be used to make dramatic improvements to people’s lives. And how psychology PhDs can be effective leaders in such non-academic roles.

User Centred Design

A UCD classic, published in 1988, needed little updating for a new edition in 2012

A few examples in the first lecture and the students were saying “this is obvious, why isn’t everything done this way?” Don Norman‘s first few pages prompt the same reaction.

It should be obvious whether a door should be pushed or pulled, without labels. A first project: simply observe people as they approach different doors around campus.

Next, ask someone to open an unfamiliar microwave, or turn on the light on the left, or mute this mobile phone. Ask them to ‘think aloud’ as they approach the task, then watch without interrupting.

No RCT, control group, or fancy statistics required. Watch just a few users before–if you can–redesigning the object or system. Then do it again, and again. But you need real users–old, young, left-handed, neurodiverse, not speaking your language, perhaps disabled…

It can be frustrating. You quickly start noting poor design everywhere: why must I type my email in two different places? Why can’t I dislike all the options? Why do I need my glasses to figure out how to turn on the upper shower? …lots of scope for your psychology students to improve the world!

Landing Aircraft Safely

Ron Cumming, my father, was an aeronautical engineer researching crashes at the most dangerous moment in flying–landing. He decided it was largely a perceptual problem–the pilot perhaps needing to land with no visible horizon and only a single line of lights down one side of the runway.

Ron spent 1960 with human factors guru Paul Fitts at the University of Michigan studying psychology. Then he and psychologist Ross Day introduced human factors (ergonomics) to Australia.

Ron and colleagues developed T-VASIS in the 1960s. It was a simple array of lights each side of the runway that indicated to the pilot whether the aircraft was on the ideal glidepath, or needed to fly up or down a little. It was installed in many countries, and some airfields still offer it, despite the widespread use of modern radar systems.

T-VASIS: The view from the cockpit, coming in to land

Leading in Road Safety Was Just the Start

After leading the world with compulsory seat belts, in 1976 Victoria was also first with random breath tests of drivers, with .05 as the alcohol limit. Ron and Ross, as the two psychology professors at Monash University, set up the Monash University Accident Research Centre. The good work of MUARC continues.

In 1987 Victoria set up VicHealth as the world’s first health promotion foundation under legislation claimed to set the standard for international best practice by banning tobacco advertising and diverting those advertising dollars to fund anti-smoking campaigns and buy out tobacco sponsorship of sport and the arts.

The Slip-Slop-Slap campaign: Slip on a shirt, Slop on sunscreen, Slap on a hat,

In 1988 VicHealth funded launch of SunSmart, with the aim of changing behaviour to increase protection against UV in sunlight and thus reduce skin cancer. For example, with its Slip-Slop-Slap campaign.

These and other public behaviour change programs have saved numerous lives, averted much suffering, and reduced health costs. Psychologists continue to play prominent roles in all these programs.

Carla Groom Diagnosed as Autistic at 44

Dr Carla Groom, the then Head of Human Centred Design Science, U. K. Department of Work and Pensions was interviewed about her later-in-life formal diagnosis as autistic, at age 44. Again she is fascinating, here in recounting how the diagnosis explained for her so much about herself and how she worked.

She tells of the strategies she uses to support other neurodiverse people to be their most effective and, more generally, how she builds diverse–usually multi-disciplinary–teams, which tend to be better at problem solving.

Training PhDs to Be Effective in Non-Academic UCD and Nudge Work

Ten Years Applying Psychological Science Inside the U.K. Government offers lessons for how postgraduate education can be broadened to prepare PhDs to work and lead effectively beyond the academy. Including: study qualitative research methods, learn to write for a general audience as well as for academic journals, work in multi-disciplinary groups, and undertake messy real-world projects.

Such as helping pilots to land safely, or nudging everyone to wear a wide-brimmed hat and apply sunscreen.

Geoff

Brunch With Petra at the Museum

Petra Vaiglova is now Senior Lecturer in Archaeological Science at the Australian National University in Canberra. I posted here about meeting her for the first time, when she visited The University of Melbourne a few years back.

Geoff, Petra, and Stephen. Behind is a very early model Holden, manufactured in Australia, towing a caravan also of about 1950.

Since then she has published two especially notable open-access articles, the first being How can we improve statistical training in archaeological science?

Improving statistical training in archaeology science

Here’s the graphical abstract, the great work of Kathryn Killackey:

See the P.P.S. below for the full abstract.

Teeth, and ritual feasting a long long time ago

The second article, also open-access, is Transport of animals underpinned ritual feasting at the onset of the Neolithic in southwestern Asia, which reports a highly innovative study, led by Petra, of teeth from an Early Neolithic site in Iran.

I was recently in Canberra and, happily, could catch up with Petra and her partner, Stephen, for brunch at the National Museum cafe. I enjoyed a very good brunch, with animated and highly interesting discussion.

Geoff

P.S. Petra also very kindly made the trek from Canberra to join a rcent party marking my 80th birthday.

Yikes!

P.P.S. The abstract of Petra’s statistics article:

  • Raising the standard for statistical training in archaeology will improve the breadth and depth of archaeological science.
  • Improving statistical training can start by discussing five fundamental statistical concepts that archaeologists do not talk about enough.
  • Supervisors can help make statistical training more effective by advocating for statistical reform and Open Science.

The aim of this paper is to shine light on fundamental statistical concepts that archaeologists do not talk about enough. I argue that more deliberate discussion of these statistical ‘elephants in the room’ can have a positive impact on improving statistical training and on steering us away from perpetuation of poor research practices.

1) Statistical thinking should come first. This will help us break down some of the stigma around numbers and statistics, and set us up for building analytical frameworks that will provide the most informative answers to our research questions.

2) Descriptive and inferential statistics have different interpretative potential. This will clarify how we can move from using tools that only allow us to talk about our studied samples to using tools that enable us to draw inferences about the underlying populations from which the samples derived.

3) p values can be extremely variable. This will help spread awareness about the misuses and misconceptions of Null Hypothesis Significance Testing (NHST) and demonstrate the dangers of using significance thresholds to interpret data.

4) Statistical precision is not the same as measurement precision. This will bring attention to the many different types of uncertainties that are built into archaeological datasets (e.g., statistical precision, instrument measurement error, natural variation),.Recognising this is key for drawing reliable inferences from our data.

5) Meta-analyses and forest plots can be useful for synthesising previous research. This will help spread awareness about the benefit of meta-analyses for creating evidence-driven summaries of previous findings.

The discussion draws on examples from isotope archaeologybioarchaeology, and organic residue analysis to illustrate how switching from a reliance on significance testing to a reliance on effect sizes can improve methodological rigour and the representativeness of our findings. The paper ends with a discussion of the roles and responsibilities of supervisors for creating an effective learning environment for statistical training. This includes, but is not limited to, acknowledging the problems of NHST and advocating for adherence to Open Science principles. Ultimately, the changes suggested in this paper will help us raise discipline-wide standards for quantitative training and improve both the breadth and the depth of archaeological research.

Thankyou Jiangang! A Ten-Year Journey to Significance Roulette

Jiangang Xia is an enterprising professor at the University of Nebraska who, among many other things, teaches into China. He alerted me some years ago to the difficulty his students in China had accessing my videos because YouTube was blocked for them. So I mounted the videos at this OSF site.

As another enterprising step, Jiangang has recently been making a number of posts to LinkedIn. Below is one of these:

From a LinkedIn post by Jiangang Xia

Rethinking Quantitative Reasoning in Educational Research (4)

Jiangang Xia

Jiangang Xia

Associate Professor of Educational Administration at University of Nebraska–Lincoln

January 13, 2026

Geoff Cumming and the New Statistics: Estimation as a Way of Thinking

Some years only reveal their importance in hindsight. For me—and for statistics reform—2014 was one of those years.

That was the year I had just begun my academic career at the University of Nebraska–Lincoln. It was also the year Rex Kline visited UNL and delivered his keynote, “Hello, Statistics Reform.” (for Kline’s talk, please see my Post 2 for more details). At the time, I had no idea how much that visit would eventually shape my thinking, teaching, and research.

What I did not know then—and only came to appreciate much later—is that 2014 was also the year Geoff Cumming delivered his workshop “The New Statistics: Estimation and Research Integrity” at the APS Annual Convention in San Francisco.

I was completely unaware of that workshop.

In fact, I would remain largely unaware of Cumming’s work for several more years—even though it quietly passed through my academic life more than once.

Missed encounters (seen only in hindsight)

While preparing this post, I did something simple: I searched Geoff Cumming’s name in my old emails. That’s when I realized how often his work had crossed my path without fully registering.

In 2016, I co-chaired a dissertation in which the student cited Cumming’s influential article “The New Statistics: Why and How” (2014). At the time, neither of us fully grasped what the article was really asking us to reconsider. Although reform language appeared, the dissertation still relied on phrases like “marginally significant” for p values above .05—an indication that NHST logic remained firmly in place. Estimation had been encountered, but not yet learned as a way of thinking.

In 2017, I received an email from Routledge inviting faculty to request a desk copy of Introduction to the New Statistics. I didn’t request it. Another missed opportunity—one I only recognize now, looking backward.

These weren’t personal oversights so much as reflections of how deeply NHST was normalized in our training. Reform ideas were present, but the infrastructure for learning them—how to teach them and how to use them—was still thin.

From critique to alternative

It wasn’t until 2019, through Rex Kline’s work, that the larger picture finally came into focus for me. From there, I traced the reform movement backward—and Geoff Cumming’s role became unmistakable.

Cumming and his colleagues were not simply extending Jacob Cohen’s critique of null hypothesis significance testing. Cohen had already shown, powerfully, why NHST was flawed and had pointed toward alternatives. He also played a central role in the APA Task Force on Statistical Inference, whose report was released in 1999.

Unfortunately, Cohen passed away in 1998, before that work could be fully carried forward. When the APA’s 5th edition was published in 2001, many of the reform ideas were only partially adopted, and everyday research norms remained largely unchanged.

What Cumming did next was different.

Through decades of scholarship—culminating in The New Statistics—he articulated a coherent alternative centered on estimation rather than binary decisions, emphasizing effect sizes, confidence intervals, precision, uncertainty, and cumulative evidence. By the New Statistics, Cumming meant an estimation-centered approach to inference—focusing on effect sizes, confidence intervals, and uncertainty, and on combining evidence across studies through meta-analysis—rather than making binary decisions based on p values alone.

From alternative to institutionalization and teaching

That work did not remain theoretical.

In 2008, Cumming was invited to join the small working group responsible for statistical reporting standards in the APA’s 6th edition, where he was a driving force behind the requirement that effect sizes and confidence intervals be reported for every research question—helping move estimation from an optional supplement to a core reporting standard.

Just as importantly, Cumming took the initiative to teach this alternative—by writing new articles (2014), new textbooks (2013, 2017, 2024), offering workshops (2014 APS), and developing demonstrations (ESCI) aimed at broad audiences, not just methodologists.

Why experience matters in teaching reform

Along the way, Cumming recognized something more fundamental: logical arguments alone were not enough.

As he later reflected, perhaps NHST had become “the researcher’s heroin—an addiction impervious to reason.”

If that was true, persuasion would require more than explanation. It would require experience.

This insight shaped his teaching. Rather than debating p values in the abstract, Cumming showed researchers—often viscerally—how unstable they are through demonstrations such as the dance of the p values, p intervals, and significance roulette. The goal was not just to convince the mind, but to engage the gut—to help researchers feel uncertainty rather than deny it.

A decade later, the field itself began to catch up. In 2024, Cumming’s 2014 article “The New Statistics: Why and How” received SAGE’s 10-Year Impact Award, recognizing research whose influence endures well beyond the standard citation window. The award was a reminder that reform ideas are often understood slowly—resisted early, adopted unevenly, and acknowledged only after they have quietly reshaped teaching and research norms.

Full circle: from missed encounters to transformed practices

What changed for me after 2019 was not just what I read—but how I taught, mentored, and conducted research.

In 2022, I shared Cumming’s 2014 article with my doctoral advisee Amanda. Her dissertation became the first I supervised to fully abandon significance language, adopting estimation-based interpretation throughout. That same year, a manuscript my student Cailen and I submitted—published in 2023—explicitly drew on Cohen, Kline, Cumming, and the ASA (2016) statement. It was my first research article grounded fully in estimation thinking.

And in a quiet but meaningful full circle, when I needed materials for my 2024 summer teaching in China, Geoff Cumming himself shared his 2014 APS workshop videos with me—materials I have since used both internationally and in my quantitative methods courses at UNL.

Looking back now, the reform was happening all around me in 2014.

I just wasn’t ready to see it.

Perhaps that is how intellectual change often works—not as a single revelation, but as a series of missed encounters that eventually align. And perhaps genuine reform requires not only better arguments, but better ways of helping researchers experience uncertainty.

For readers who want to explore further

(All materials shared with permission; enormous credit to Geoff Cumming, Bradley Dean, and Robert Calin-Jageman.)

Dear colleagues,

  • When did you first encounter Geoff Cumming’s work—if at all?
  • Have you seen estimation treated as an add-on, rather than a way of thinking?
  • What important ideas did you meet early, but only understand much later?

#Cumming #NewStatistics #EstimationThinking #QuantitativeMethods #EducationalResearch

Thank you Jiangang!

Geoff

A Statistics Textbook for the AI Era

Miodrag Lovrić is an enormously energetic statistician and educator. He persuaded 700 scholars from 110 countries to contribute to the massive four-volume second edition of the International Encyclopedia of Statistical Science (Springer, 2025).

Now he is close to completing Statistical Thinking for the AI Era, which comprises Vol. 1: Foundations and Inference, and Vol. 2: Advanced Methods–each with about 15 chapters and 520 pages. The chapters I have seen indicate it’s a terrific textbook.

Miodrag has lived and worked on at least four continents and takes a wholeheartedly global focus in both those major works. The textbook starts with Forewords from many continents. Miodrag generously invited me to contribute a foreword from Australia–which appears below. As you see, I slipped in mention that I’m an Antarctic tragic, even if it’s a big claim that I speak also from that seventh continent!

Foreword From Australia and Antarctica

I’ve been an Antarctic tragic since I was a boy. I’ve participated in citizen science during two voyages south. I’ve observed glacial retreat and penguin colonies dying. I read about giant datasets analysed by AI-assisted statistical modelling of changes to sea ice, weather, and much else including terrifying tipping points that will determine catastrophic changes our children must endure.

I’ve given a well-received statistics talk to researchers at the Institute for Marine and Antarctic Studies in a building on the wharf where Antarctic supply ships berth. All around us were whale skulls, ancient sledges, and other memorabilia. I feel I can claim at least a little justification for offering this piece on behalf of the seventh continent, as well as speaking from Australia.

Statistics is about communication and therefore the province of psychologists. Researchers publish representations of their results and readers, ranging from researchers to politicians, policymakers, and ordinary people, draw conclusions from those representations. However, vast industries misrepresent research results as they seek to persuade us to support autocrats, eat unhealthy food, buy things we don’t want that trash the planet, and burn fossil fuels to destroy the chance our children will inherit a liveable world.

My students and I studied people’s misconceptions of basic statistical concepts. That was statistical cognition, which has now developed hugely into metascience—a wonderfully interdisciplinary field that investigates how research is done, and should be done to be more trustworthy. For example, it studies how paper mills use AI tools to generate fake manuscripts that desperate researchers buy then submit to journals. Quality journals use AI tools and much editorial expertise to try to weed out fakes before these criminally pollute the research literature—much of it on life-and-death issues.

Around 2014 Open Science emerged: replication is central; this requires meta-analysis to synthesise results; this requires point and interval estimates, and statistical significance is irrelevant and often damaging. My statistics textbooks advocate estimation and meta-analysis. The intro book, with Bob Calin-Jageman, integrates estimation and Open Science all through. The second edition (Routledge, 2024) includes Bob’s wonderful open source software that’s ideal for beginners and researchers aiming for good estimation and Open Science practices. At thenewstatistics.com is more information, also the significance roulette simulation, which dramatizes a highly misleading feature of p values that few researchers appreciate. The good news is that these new approaches are found accessible by students and are a delight to teach.

I conclude that everyone should have a basic understanding of evidence and statistics. This book is broad in scope, well-informed, and future-focussed. For example, it explains traditional, Bayesian, and randomisation frameworks for estimation, any of which can support Open Science. As I read, I’m often applauding the approaches Professor Lovrić has chosen. With this book he is making an enormous contribution to statistics education around the world.

Geoff

John Self, AI Pioneer, Chats With ChatGPT

A Chat with ChatGPT by a veteran AI researcher

John Self entered the field of AI in the early 1970s. He’s a distinguished scholar who can claim to have introduced the idea of user model in his 1974 article (while visiting The University of Melbourne). He was writing in the context of AI in Education, a field in which he was a pioneer and long-time leader. Released in 2005, his last AI book is humane and broad, on open access and an excellent read: Whoever Said Computers Would be Intelligent 

John generously hosted what was for me an important sabbatical, in Lancaster in 1988. For a couple of decades back then my research was as a psychologist in AIEd. Many in that field (not John, and not some of the other leading lights) were IT tragics with a view that human learning was little more than turning on the tap to fill the bucket. John’s landmark contribution in 1974 was to recognise that any IT (‘intelligent tutor’, the pretentious term back then) that individualised its response to a student must contain a model of that student. The crudest might be merely a note of where that student was up to in the book. Typically it would comprise a record of previous student work, correct responses and errors, and earlier comments by the IT.

Early Intelligent Tutors (ITs)

Many of those early ITs supported learners to achieve impressive gains on tests, especially in science and computing fields. Interactions resembled those in ‘direct instruction’ classrooms. Direct instruction is highly structured and is coming back into vogue in many countries, often for early reading and numeracy. However, then as now, many students find the interactions stultifying. Grit your teeth and use an IT to quickly get up to speed with LISP! But a fulfilling education, perhaps not 🙁

Learner Models in ITs

My role in AIEd was often to be the maverick critic, despairing of the poverty of typical learner models. I and colleagues studied transcripts of tutoring interactions of expert teachers with individual learners when given occasional chances to interact. Think of a teacher wandering in a class working as individuals. The teacher has a brief interaction with an individual, initiated by a student requesting help, or the teacher walking by and choosing to interrupt.

Human Tutoring and Learner Models

Not surprisingly, the interactions were often brief, just a single Q&A in either direction, or little more. But they were highly diverse. The Q might be about motivation, feelings, seemingly irrelevant things that affect the learner’s work and thinking. Brief explanations can of course be valuable, but possibly the most valuable comments were often much higher-level, about strategy, or motivation. “Inspiring” is possibly the most valuable teacher ability, as hopefully we all remember. Human teachers have learner models for their individual learners. These models are likely fragmentary in many respects, but they are, most importantly, highly diverse. Achieving this richness was the enormous challenge for IT researchers seeking to model good tutoring. It remains a core challenge for any AI intended to be used by a person.

Rainbow over Kisdon in Swaledale: From the home page of Saunterings

John the Fell Runner and Hill Walker

For decades John was an immensely fit fell runner, spending hours and days in the hills of North-West England. He stopped running in 2017, then from 2018 has been posting online reflections and great photos from his Saunterings in the hills, dales, and moors near Lancaster and surrounds.

ChatGPT’s User Model

John started by asking “What do you think of Saunterings by John Self?” I’m guessing John had front of mind trying to diagnose ChatGPT’s model of its user–John. What did it know about, what did it assume about John? John gives his own commentary, commenting on the AI responses and telling us a little of his own thinking. Lots of fascinating stuff there, especially as John picks apart what seems to be underlying the conversation.

ChatGPT adopts a chatty, deferential, friendly, explanatory style, quick to apologise and explain its own errors. John could of course ask it to adopt a different perspective and style–it immediately did so, and quite convincingly.

John’s Insights

John’s first overall reaction was astonishment that ChatGPT could do so well, and so blindingly fast. Then follows pure gold as John extends his Q&A to investigate various thoughts about how the IT works, what assumptions it makes, how it formulates opinions, and where its boundaries lie. To what extent does it build a model of John? To what extent is it merely integrating the results from a huge number of searches, or is it reasoning about these? John eventually concludes that he cannot trust ChatGPT and cannot follow its reasoning–because there isn’t any. I won’t try to summarise: you need to read John’s final paragraphs to get the rich story.

Extremely well-informed gold!

Geoff

Vale Michael Kubovy (1940-2025), Professor of Patterns

I don’t think I ever met Michael, but have long known his Gestalt perception work. I now discover he did so much more, especially as a pioneer in data analysis. He and I would have agreed on many, many things.

This post is courtesy Alex Holcombe, who wrote:

A tribute to Michael Kubovy

You can browse Michael’s books on perception here, and his dazzling ‘Psychology of Perspective and Renaissance Arthere.

During WW2 he fled with his parents from France to Portugal, then eventually Israel, where he completed his university education with cog psy royalty: his master’s with Daniel Kahneman, who initially hired him as a laboratory assistant after a chance encounter at a corner shop, and his doctorate with Amos Tversky, whom he met as his commander in the Israeli reserves. 

Among his notable inventions was an auditory analogue to the random-dot stereogram. This allowed a listener to hear a hidden melody with their “third ear” that was entirely undetectable by either ear alone.

Considering data analysis, Michael took delight in the pleasures of taking data seriously. This meant finding a way to visualise and explore data, a key interest of statistician John Tukey. He loved the early software Data Desk, which even allowed you to use sliders to interact with a data visualisation!

Michael lamented reliance on null hypothesis significance testing (NHST), a major cause of the replication crisis. He left, at his death, a sadly unfinished book designed to provide a solution: Use visualisation to evaluate quantitative models. Part of his description:

“We have tools to see if the residuals from the model are normally distributed. And so what we do is we teach the students to use certain tools in a routine way to see if data deviate in any important way from the model that they’re proposing. So we use graphics a lot and we show by example. And because this is a textbook based on R, and there will be R code all over the place, they will have examples of good practice. We’re not going to preach, but we’re going to give examples of what to do. And we hope that by osmosis and by practicing with our examples that we give in the text and on our website, we hope that people will acquire best practices in their data analysis work.”

Vale Michael Kubovy.

Geoff

PS It’s worth reading Alex’s full tribute 🙂

Stop Fooling Yourself – paper by Richard Born

eNeuro has a great new paper by Richard Born (Harvard) about how to diagnose and avoid confirmation bias (Born 2024). The full text is here: https://www.eneuro.org/content/11/10/ENEURO.0415-24.2024

This paper is part of a series on Improving Your Neuroscience that I (R.C-J.) am helping to organize (Calin-Jageman 2024). You can find a list of all papers as they appear in this series here: https://www.eneuro.org/collection/improving-your-neuroscience

I strongly recommend this lovely paper – it is full of fascinating examples and references; it is the type of paper you’ll want to assign to all your incoming trainees.

References

Born, Richard T. 2024. “Stop Fooling Yourself! (Diagnosing and Treating Confirmation Bias).”eNeuro 11 (10). https://doi.org/10.1523/ENEURO.0415-24.2024.

Calin-Jageman, Robert J. 2024. “New eNeuro Series: Improving Your Neuroscience.”eNeuro 11 (3). https://doi.org/10.1523/ENEURO.0048-24.2024.

The ideal statistics curriculum?

This week I (Bob) was part of a workshop on the future of neuroscience education. I got to be part of a rockstar panel with Tari Tan (Harvard), Monica Linden (Brown), and Rosalind Segal (Harvard). The event was organized by Lique Coolen (Kent State).

Tari spoke about curricular goals for neuroscience, reviewing an SFN initiative to define core competencies for trainees at the undergraduate and graduate level. I didn’t know about this! Tari gave a great overview; it’s well worth checking out: https://www.sfn.org/careers/higher-education-and-training/core-competencies

Monica discussed the role of generative AI in the future of neuroscience education and gave lots of thoughtful examples and resources.

Rosalind discussed developing an internship program for the neuroscience *PhD* program at Harvard! Very thougtful and interesting way for grad students to explore the diverse career paths that can come out of doctoral training — but also raised lots of interesting issues (a stated goal of internships that don’t impede research seemed hard to balance; students seemed to express worries about losing faculty support if they expressed ‘wrong’ career goals… lots to think about!).

My talk was about the ideal statistics curriculum. My short answer was “There isn’t one!” — neuro is too diverse and there is not only one ‘right’ way to do statistical inference. Still, the field of neuroscience shows some evident difficulties making valid claims from data (on that note, check out Chen et al., 2024), and so at least thinking about some broad ideals for training seems like a useful exercise. I came up with these guiding principles:

  • Estimation thinking at the forefront
  • Teach using simulations — for exploration at first, but eventually for planning experiments before they are conducted
  • Quickly graduate from toy examples to real, complex data sets and projects that require critical thinking about Multiplicity and Interdependence, and teach robustness checking to help students validate their approaches.
  • Integrate Open Science throughout
  • Ensure training is not just about statistics, but about the ‘neglected factors’, including good design and measurement — strength of evidence and quality of research are about so much more than just the statistics generated, and improvements are about so much more than sample size.

My talk and some resources on each of these topics are here: https://osf.io/muy6u/wiki/Workshops%20-%20SFN%202024/

We’re collating all the resources from all the speakers; I’ll post a link when I have that.

1635500 {1635500:JY7HWM5Q},{1635500:QX7U9XKT} 1 apa 50 default 2638 https://thenewstatistics.com/itns/wp-content/plugins/zotpress/
%7B%22status%22%3A%22success%22%2C%22updateneeded%22%3Afalse%2C%22instance%22%3Afalse%2C%22meta%22%3A%7B%22request_last%22%3A0%2C%22request_next%22%3A0%2C%22used_cache%22%3Atrue%7D%2C%22data%22%3A%5B%7B%22key%22%3A%22JY7HWM5Q%22%2C%22library%22%3A%7B%22id%22%3A1635500%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Chen%22%2C%22parsedDate%22%3A%222022-05-01%22%2C%22numChildren%22%3A2%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BChen%2C%20D.%20%282022%29.%20%26lt%3Bi%26gt%3BSTATISTICAL%20PRACTICE%20IN%20PRECLINICAL%20NEUROSCIENCES%3A%20IMPLICATIONS%20FOR%20SUCCESSFUL%20TRANSLATION%20OF%20RESEARCH%20EVIDENCE%20FROM%20ANIMALS%20TO%20HUMANS%20Committee%20Member%26lt%3B%5C%2Fi%26gt%3B.%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22book%22%2C%22title%22%3A%22STATISTICAL%20PRACTICE%20IN%20PRECLINICAL%20NEUROSCIENCES%3A%20IMPLICATIONS%20FOR%20SUCCESSFUL%20TRANSLATION%20OF%20RESEARCH%20EVIDENCE%20FROM%20ANIMALS%20TO%20HUMANS%20Committee%20Member%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Darin%22%2C%22lastName%22%3A%22Chen%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22date%22%3A%222022-05-01%22%2C%22originalDate%22%3A%22%22%2C%22originalPublisher%22%3A%22%22%2C%22originalPlace%22%3A%22%22%2C%22format%22%3A%22%22%2C%22ISBN%22%3A%22%22%2C%22DOI%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222024-10-08T15%3A15%3A56Z%22%7D%7D%2C%7B%22key%22%3A%22QX7U9XKT%22%2C%22library%22%3A%7B%22id%22%3A1635500%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Calin-Jageman%20and%20Cumming%22%2C%22parsedDate%22%3A%222019%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BCalin-Jageman%2C%20R.%20J.%2C%20%26amp%3B%20Cumming%2C%20G.%20%282019%29.%20Estimation%20for%20Better%20Inference%20in%20Neuroscience.%20%26lt%3Bi%26gt%3BEneuro%26lt%3B%5C%2Fi%26gt%3B%2C%20%26lt%3Bi%26gt%3B6%26lt%3B%5C%2Fi%26gt%3B%284%29%2C%20ENEURO.0205-19.2019.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.1523%5C%2FENEURO.0205-19.2019%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.1523%5C%2FENEURO.0205-19.2019%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22journalArticle%22%2C%22title%22%3A%22Estimation%20for%20Better%20Inference%20in%20Neuroscience%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Robert%20J.%22%2C%22lastName%22%3A%22Calin-Jageman%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Geoff%22%2C%22lastName%22%3A%22Cumming%22%7D%5D%2C%22abstractNote%22%3A%22The%20estimation%20approach%20to%20inference%20emphasizes%20reporting%20effect%20sizes%20with%20expressions%20of%20uncertainty%20%28interval%20estimates%29.%20In%20this%20perspective%20we%20explain%20the%20estimation%20approach%20and%20describe%20how%20it%20can%20help%20nudge%20neuroscientists%20toward%20a%20more%20productive%20research%20cycle%20by%20fostering%20better%20planning%2C%20more%20thoughtful%20interpretation%2C%20and%20more%20balanced%20evaluation%20of%20evidence.%22%2C%22date%22%3A%2207%5C%2F2019%22%2C%22section%22%3A%22%22%2C%22partNumber%22%3A%22%22%2C%22partTitle%22%3A%22%22%2C%22DOI%22%3A%2210.1523%5C%2FENEURO.0205-19.2019%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Fwww.eneuro.org%5C%2Flookup%5C%2Fdoi%5C%2F10.1523%5C%2FENEURO.0205-19.2019%22%2C%22PMID%22%3A%22%22%2C%22PMCID%22%3A%22%22%2C%22ISSN%22%3A%222373-2822%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%226ECZTSSQ%22%2C%22M4C95AAC%22%5D%2C%22dateModified%22%3A%222023-01-17T22%3A06%3A18Z%22%7D%7D%5D%7D
Chen, D. (2022). STATISTICAL PRACTICE IN PRECLINICAL NEUROSCIENCES: IMPLICATIONS FOR SUCCESSFUL TRANSLATION OF RESEARCH EVIDENCE FROM ANIMALS TO HUMANS Committee Member.
Calin-Jageman, R. J., & Cumming, G. (2019). Estimation for Better Inference in Neuroscience. Eneuro, 6(4), ENEURO.0205-19.2019. https://doi.org/10.1523/ENEURO.0205-19.2019

‘The New Statistics’ (2013) Wins Sage 10-Year Impact Award

The New Statistics: Why and How (abstract below) explained the advantages of moving on from NHST to the new statistics (estimation and meta-analysis) and the need for better practices to improve research integrity. I’m delighted that an award from Sage indicates the article seems to be helping researchers improve what they do. Next: Can ITNS2 help the next generation do even better?

The article appeared online in late 2013, so was considered when Sage examined the citation numbers of all articles appearing in any of the 400+ journals Sage published back in 2013. It was one of the top three most cited, so has been given a Sage 10-Year Impact Award. Sage’s announcement is here. Sage has just published a blog post about it with headline:

Now for the abstract:

The article was commissioned by Eric Eich, then editor-in-chief of Psychological Science, to appear immediately following his famous editorial Business Not As Usual. This opened the Journal’s first issue of 2014 and announced sweeping changes in the journal’s submission requirements, which, for many psychologists, marked the arrival of Open Science.

Interview

Sage’s blog post includes an email interview with me. Here’s a brief summary:

What was it in your own background that led to your article?

When I was a teenager my father gave me a simple explanation of significance testing. I said something like “That’s weird, sort of backwards. And why .05?” He replied “I agree, but that’s the way we do it.”

Over decades of teaching I became ever more dissatisfied with NHST, and focused ever more on confidence intervals (CIs).

Was there an article that had a particularly strong influence on you?

Frank Schmidt (1996) wrote: “It is now possible to use meta-analysis to show that reliance on significance testing retards the development of cumulative knowledge.” A revelation!

What did Schmidt’s article lead to?

About 2003 I started using an Excel forest plot to give a simple explanation of meta-analysis in my intro course. I was delighted: Students told me it just made sense. Of course, for meta-analysis you need a CI from each study, while p values are irrelevant, even misleading.

      Figure: Dances of means, confidence intervals, and p values.

In 2009 I uploaded a video of the dance of the p values. I became passionate about advocating the new statistics (estimation and meta-analysis). I wrote Understanding The New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis (UTNS, 2012).

What was happening in psychology at about that time?

Ioannidis (2005) explained how reliance on NHST was a major cause of the replication crisis. Largely in response to that crisis, Open Science arrived—perhaps the most important advance in how science is done for a very long time.

Eric Eich’s famous editorial Business Not As Usual in the January 2014 issue of Psychological Science marked the arrival of Open Science in psychology. Months earlier Eric had invited me to write a tutorial article to support the changes he wanted. This was The New Statistics: Why and How and was published immediately following his editorial.

What has been the reception of the article?

Mainly very positive. Some have felt I went too far in advising that in most cases it’s better not to use NHST at all. Some Bayesians have been unhappy with the focus on confidence intervals.

Revisiting that article, what would you have done differently?

I used the term ‘research integrity’, but ‘Open Science’ was coming into use and I soon realized that was way better. Reading the article today, for ‘research integrity’ read ‘Open Science’.

Otherwise, I think the article has held up well, including all 25 guidelines in Table 1.

What has happened since

Psychological Science has continued to lead in the adoption of Open Science practices.

Meta-science, also known as meta-research, has emerged and now thrives as a highly multi-disciplinary field. It applies the scientific method to improve that method—wonderful!

What have you been doing since?

I teamed with Robert Calin-Jageman to write the first intro statistics textbook based on the new statistics and with Open Science all through. The second edition has just come out: Introduction to The New Statistics: Estimation, Open Science, and Beyond, 2nd edition (ITNS2, 2024). It has much improved software, as we explain in Calin-Jageman & Cumming (2024), which is on open access.

We believe this book can sweep the world—we’ll see! To read the Preface and Chapter 1 go to www.thenewstatistics.com. In the second para is a link to the book’s Amazon site. Click ‘Read sample’.  

References

Calin-Jageman, R., & Geoff Cumming, G. (2024). From significance testing to estimation and Open Science: How esci can help. International Journal of Psychology,     https://doi.org/10.1002/ijop.13132

Cumming, G. (2012). The New Statistics: Effect sizes, confidence intervals, and meta-analysis. New York: Routledge. 

Cumming, G. (2014) The new statistics: Why and how. Psychological Science. 25(1), 7-29. https://doi.org/10.1177/0956797613504966

Cumming, G., & Calin-Jageman, R. (2024). Introduction to The New Statistics: Estimation, Open Science, & Beyond, 2nd edition. New York: Routledge.

Eich, E. (2014) Business not as usual. Psychological Science, 25(1), 3–6. https://doi.org/10.1177/0956797613512465

Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine 2: e124. https://doi.org/10.1371/journal.pmed.0020124

Schmidt, F. L. (1996). Statistical significance testing and cumulative knowledge in psychology: Implications for training of researchers. Psychological Methods, 1(2), 115-129. https://doi.org./10.1037/1082-989X.1.2.115