The p Value Casino Is Open–For Significance Roulette!

Excel 2003, the best version ever, was enshittified by MicroSoft in the 2007 version, which was way slower and dropped many wonderful animation facilities :-(. Even vast efforts would not get my great Significance Roulette simulation running in the new version.

Two videos use my Excel 2003 version to explain: https://tiny.cc/SigRoulette1 and https://tiny.cc/SigRoulette2 or simply search at YouTube for ‘significance roulette’.

Now Bradley Dean has built Significance Roulette to run in your browser–as part of esci web.

Significance Roulette, after an initial study gives p = .01

It has taken me close to 20 years to build the following argument that leads to significance roulette, and which is summarised in the first half of our open access article Calin-Jageman & Cumming, 2024

  • For approaching a century numerous distinguished scholars, including philosophers of science, statisticians, and psychologists, have published cogent critiques of p values, significance testing, and how researchers across science use these.
  • Even so, a large proportion of researchers, teachers, journals, and granting bodies use null hypothesis significance testing (NHST)–often based on p < .05 or p < .01–as the standard for concluding whether or not an effect exists, whether or not a result is large or important. Despite this logic being wrong-headed in so many ways!
  • Devotion to NHST and p <.05 resembles an addiction–the researcher’s heroin. Rational argument is not sufficient to shake the addiction. Could a dramatic demonstration, perhaps persuading via the gut rather than the brain, shake this addiction?
  • A striking but little-known feature of p values is that they are highly unreliable–their sampling variability is astonishingly large. Replicate a study, exactly the same but with a new random sample, and expect to obtain replication p that can take just about any value between 0 and 1!
  • Jerry Lai in his lovely PhD studies took three converging approaches to find that a large proportion of published researchers in psychology, medicine, and statistics severely under-estimate the amount of variability in the p value with replication.
  • My first demonstration of p value variability was the dance of the p values, the first video of which dates from 2009. Search YouTube for ‘dance of the p values’ to find several videos. You can also play with the dances in esci web.
  • My second approach arose from study of the probability distribution of replication p, the p value obtained in a replication. My highly-cited 2008 article has details.
  • I define the p interval as the 80% prediction interval for replication p. It’s astonishingly long! For example, if an initial study obtains p = .05, the p interval is (.0002, .65), meaning an 80% chance of p within that interval and fully a 10% chance it falls below .0002 and 10% above .65. After p = .01 the interval is (.00001, .41). After initial p = .001 (*** highly statistically significant) there is fully a 1 in 6 chance a replication does not even achieve p < .05! After p = .20 ns there is a 1 in 3 chance a replication finds p < .05!
  • I take the probability distribution of replication p following initial p = .01 and divide the area under the curve, which represents probability, into 38 equal areas. I label each area with the p value in the centre of the interval. I have 38 p values, a few large and many small and very small, which accurately represent that probability distribution.
  • I scatter those 38 p values randomly around the 38 bins of a roulette wheel. Simply click to spin, wait for a few moments, and see the replication p you might have obtained from a replication. Much faster and cheaper than the hassle of raising a grant, recruiting participants, hassling with ethics approval… and collecting data!
  • At Significance Roulette in esci web you can click between initial p of .05 and .01. With initial p = .01, replication p values tend a little smaller, of course. But the striking thing is how widely spread the p values are! The distributions are pictured to left of the wheel. For example, switch from .05 to .01 and note a slightly smaller number of deep blue, deep trombone sound, despairing figures for p > .1 ns and slightly more bright red, triumphant trumpet blast, elated figures for p < .001 ***. Click SPIN, and note your quickened heart beat, sweaty palms, and that you are holding your breath–will you be despairing or elated–and in only a few seconds you’ll know!

Will this approach to tackling addiction via the gut be more effective than the decades of argument addressed at the cortex?

Best of luck at the p Value Casino!

Take-home messages

  • p values are unbelievably unreliable
  • Any p value could easily have been just about any other value
  • No p value deserves our trust 🙁
  • Simply don’t use p values, there are much better ways 🙂

Geoff

P.S. Enormous thanks to Bradley and Bob, who made it all happen.

John Self, AI Pioneer, Chats With ChatGPT

A Chat with ChatGPT by a veteran AI researcher

John Self entered the field of AI in the early 1970s. He’s a distinguished scholar who can claim to have introduced the idea of user model in his 1974 article (while visiting The University of Melbourne). He was writing in the context of AI in Education, a field in which he was a pioneer and long-time leader. Released in 2005, his last AI book is humane and broad, on open access and an excellent read: Whoever Said Computers Would be Intelligent 

John generously hosted what was for me an important sabbatical, in Lancaster in 1988. For a couple of decades back then my research was as a psychologist in AIEd. Many in that field (not John, and not some of the other leading lights) were IT tragics with a view that human learning was little more than turning on the tap to fill the bucket. John’s landmark contribution in 1974 was to recognise that any IT (‘intelligent tutor’, the pretentious term back then) that individualised its response to a student must contain a model of that student. The crudest might be merely a note of where that student was up to in the book. Typically it would comprise a record of previous student work, correct responses and errors, and earlier comments by the IT.

Early Intelligent Tutors (ITs)

Many of those early ITs supported learners to achieve impressive gains on tests, especially in science and computing fields. Interactions resembled those in ‘direct instruction’ classrooms. Direct instruction is highly structured and is coming back into vogue in many countries, often for early reading and numeracy. However, then as now, many students find the interactions stultifying. Grit your teeth and use an IT to quickly get up to speed with LISP! But a fulfilling education, perhaps not 🙁

Learner Models in ITs

My role in AIEd was often to be the maverick critic, despairing of the poverty of typical learner models. I and colleagues studied transcripts of tutoring interactions of expert teachers with individual learners when given occasional chances to interact. Think of a teacher wandering in a class working as individuals. The teacher has a brief interaction with an individual, initiated by a student requesting help, or the teacher walking by and choosing to interrupt.

Human Tutoring and Learner Models

Not surprisingly, the interactions were often brief, just a single Q&A in either direction, or little more. But they were highly diverse. The Q might be about motivation, feelings, seemingly irrelevant things that affect the learner’s work and thinking. Brief explanations can of course be valuable, but possibly the most valuable comments were often much higher-level, about strategy, or motivation. “Inspiring” is possibly the most valuable teacher ability, as hopefully we all remember. Human teachers have learner models for their individual learners. These models are likely fragmentary in many respects, but they are, most importantly, highly diverse. Achieving this richness was the enormous challenge for IT researchers seeking to model good tutoring. It remains a core challenge for any AI intended to be used by a person.

Rainbow over Kisdon in Swaledale: From the home page of Saunterings

John the Fell Runner and Hill Walker

For decades John was an immensely fit fell runner, spending hours and days in the hills of North-West England. He stopped running in 2017, then from 2018 has been posting online reflections and great photos from his Saunterings in the hills, dales, and moors near Lancaster and surrounds.

ChatGPT’s User Model

John started by asking “What do you think of Saunterings by John Self?” I’m guessing John had front of mind trying to diagnose ChatGPT’s model of its user–John. What did it know about, what did it assume about John? John gives his own commentary, commenting on the AI responses and telling us a little of his own thinking. Lots of fascinating stuff there, especially as John picks apart what seems to be underlying the conversation.

ChatGPT adopts a chatty, deferential, friendly, explanatory style, quick to apologise and explain its own errors. John could of course ask it to adopt a different perspective and style–it immediately did so, and quite convincingly.

John’s Insights

John’s first overall reaction was astonishment that ChatGPT could do so well, and so blindingly fast. Then follows pure gold as John extends his Q&A to investigate various thoughts about how the IT works, what assumptions it makes, how it formulates opinions, and where its boundaries lie. To what extent does it build a model of John? To what extent is it merely integrating the results from a huge number of searches, or is it reasoning about these? John eventually concludes that he cannot trust ChatGPT and cannot follow its reasoning–because there isn’t any. I won’t try to summarise: you need to read John’s final paragraphs to get the rich story.

Extremely well-informed gold!

Geoff

“The New Statistics” by INXS: Why Not?!

Thanks to Prof Dena A. Pastor, of James Madison University for this suggestion.

Whenever you hear the hit New Sensation, by Australian rock band INXS, replace “A new sensation” with “The New Statistics”. This works for me (of course it would), even if the ’80s are a bit recent for me.

Enjoy!

Geoff

What We’ll Never Know

Message just in from James Pennebaker, President of the APS, about a sadly highly important project by Tim Wilson. Tim is collecting a bunch of short punchy videos on the theme of ‘what we’ll never know’–because Trump has just cancelled our research funding.

The first four are already up at https://www.youtube.com/@TimothyWilson18

Use these to lobby your local politicians, etc. Maybe also suggest further leading researchers who might make videos. Suggestions to Tim at tdw@virginia.edu

U.S. funding cuts are also having drastic repercussions around the world–a substantial proportion of Australia’s leading researchers have U.S. funding, often as part of international collaborations. Here we’re all seeing that all-to-familiar combination of disbelief, uncertainty, despair… as highly promising programs are trashed, data destroyed, billions wasted…

Weeping,

Geoff

Vale Michael Kubovy (1940-2025), Professor of Patterns

I don’t think I ever met Michael, but have long known his Gestalt perception work. I now discover he did so much more, especially as a pioneer in data analysis. He and I would have agreed on many, many things.

This post is courtesy Alex Holcombe, who wrote:

A tribute to Michael Kubovy

You can browse Michael’s books on perception here, and his dazzling ‘Psychology of Perspective and Renaissance Arthere.

During WW2 he fled with his parents from France to Portugal, then eventually Israel, where he completed his university education with cog psy royalty: his master’s with Daniel Kahneman, who initially hired him as a laboratory assistant after a chance encounter at a corner shop, and his doctorate with Amos Tversky, whom he met as his commander in the Israeli reserves. 

Among his notable inventions was an auditory analogue to the random-dot stereogram. This allowed a listener to hear a hidden melody with their “third ear” that was entirely undetectable by either ear alone.

Considering data analysis, Michael took delight in the pleasures of taking data seriously. This meant finding a way to visualise and explore data, a key interest of statistician John Tukey. He loved the early software Data Desk, which even allowed you to use sliders to interact with a data visualisation!

Michael lamented reliance on null hypothesis significance testing (NHST), a major cause of the replication crisis. He left, at his death, a sadly unfinished book designed to provide a solution: Use visualisation to evaluate quantitative models. Part of his description:

“We have tools to see if the residuals from the model are normally distributed. And so what we do is we teach the students to use certain tools in a routine way to see if data deviate in any important way from the model that they’re proposing. So we use graphics a lot and we show by example. And because this is a textbook based on R, and there will be R code all over the place, they will have examples of good practice. We’re not going to preach, but we’re going to give examples of what to do. And we hope that by osmosis and by practicing with our examples that we give in the text and on our website, we hope that people will acquire best practices in their data analysis work.”

Vale Michael Kubovy.

Geoff

PS It’s worth reading Alex’s full tribute 🙂

Open Science: Free Zoom With the Experts Next Week

Definitely worth joining on 18 December, even if for me it’s at 6.30am. Note the first three speakers also kindly gave generous endorsements of ITNS2, at the start of the book. OS leaders, for sure.

Info and registration here. The announcement:

In 2025, SIPS will hold its 10th annual meeting! In celebration of this milestone, please join us on December 18, 2024 (see starting time in your time zone) for the warm-up event SIPS 2025 Pre-Conference DiscussionWhat went wrong? How can we do it better?, during which a few early reformers will discuss challenges and improvements in psychological science. This 90-minute virtual event is free and open to all. Our invited speakers are:

  • Simine Vazire, Professor at the University of Melbourne, co-founder of SIPS 
  • Brian Nosek, Executive Director of the Center for Open Science, co-founder of SIPS 
  • Dorothy Bishop, University of Oxford
  • Joseph Simmons, Professor at the University of Pennsylvania, co-founder of Data Colada
  • Leif Nelson, Professor at UC Berkeley, co-founder of Data Colada

Moderator: Balazs Aczel, ELTE, host of SIPS 2025 in Budapest, Hungary.  

CLICK HERE TO REGISTER You can find out more about the event at https://improvingpsych.org/sips-2025-precon-discussion (including a link to submit your questions in advance).

Vale Mac (Malcolm Macmillan): Psychologist, Historian, and True Scholar

A memorable gathering at Melbourne University recently celebrated the remarkable life and contributions of Malcolm Bruce Macmillan (Mac) (1 Jan 1929—11 Aug 2024).

Mac was one of the Monash University intro teachers in 1966 who helped me discover psychology as an absorbing science. I happily switched from physics and have been with psych and stats ever since. Thanks Mac! I’ll say a little here about two of his books.

Phineas Gage and the Tamping Iron

Numerous textbooks include brief accounts of the railway worker who, in 1848, miraculously survived—for 11 years!—a massive metal bar destroying much of his left frontal lobe. Seeking a case study to illustrate a lecture, Mac discovered how little was known of Gage and his story.

After many years and painstaking chasing of original sources round the world, in 2000 An Odd Kind of Fame appeared (MIT Press). A preview is here. It’s an enormous work, beautifully written, describing how the case has been used and misused as neuroscience has developed.

Mac’s Gage page here is a wonderful trove of (accurate!) Gage information. Read here about the 150th anniversary meeting Mac organized, which was held in Cavendish, Vermont where the accident happened. You can see here the memorial plaque unveiled during the meeting.

Freud Evaluated

Mac’s even larger exposition and critique of Freud’s work has become a classic. Arising from maybe two decades of study and writing, Freud Evaluated appeared in 1997 (MTI Press). A preview of an earlier version is here. An interview that gives a good idea of Mac’s incisive analysis (sorry!) of Freud’s ideas and methods is here.

International Society for the History of the Neurosciences

Mac was a co-founder of the Society and, for many years, an active contributor. He was also a co-founder, and co-editor 2005-2017, of the Journal of the History of the Neurosciences.

And furthermore…

I could go on. His book Snowy Campbell: Australian Pioneer Investigator of the Brain (2016). His role in developing clinical psychology in Australia… and that’s not to mention the Macmillan tartan kilt, single malts, and jazz.

Above all, a distinguished scholar, and a warm, engaging person.

Geoff

P.S. Thanks to Edith, Mac’s partner, for assistance with this post.

The Spooky Men: A Last-Minute Meditation/Entreaty From Down Under

Watch The Spooky Men’s Chorale (3 min) here. (First uploaded four years ago.)

We have followed The Spooky Men for years, with their wondrous a capella harmonies and moving lyrics. They have just sent this message:

dear world
we are the spooky men
you know that we are silly
you know that we become silliness, that silliness becomes us
we have beards, we have jokes
but also you know that underneath that we are men who hold things dear: things of goodness, warmth, strength
we believe in things, in this world
and now we stand at our darkest moment, the moment when such things as we believe face their gravest threat
years ago we asked the question: do we speak his name?
do we speak his name on stage, amongst ourselves?
the answer was, had to be no
because that would be the very thing he wanted, to add fuel to his fire
and for us, that still holds true
instead, we wrote a song
so now please forgive us, we must do the only thing we can do, even though it feels like it is not much, that there is an inferno, and we are flicking a teaspoon of water from 40 paces…
and thus we send you this song <see link at the start, above>

If the U.S. steps away from leadership in the world’s campaign against climate change, our children, grandchildren, and all future generations will pay a grievous price. In the words of Joëlle Gergis, that would be ‘An intergenerational crime against humanity’

The world holds its breath. We can’t vote. If you can, please consider.

In hope,

Geoff

ITNS2 Data Sets Are Now Available Within esci

All the data sets used in examples and exercises in the new second edition of ITNS are now available within esci in jamovi, esci in R, and soon within esci in JASP.

Here’s how to open one of those data sets in jamovi.

  1. Open jamovi with the esci module installed–meaning you can see the esci icon in the top bar, as pictured.
  2. Click three lines, top left. Panel opens.
  3. Click Open, Data Library.
  4. See 47 data files, ordered by chapter of first mention in the book.
  5. Click on your choice of data set.

In the book, ignore instructions to go to the book’s Companion Website to download a data set and save locally before opening in jamovi.

Thanks Bob, life is now easier!

Geoff