Blog

Teaching Statistics Using Web Applets–Including esci

Recognise the image? Not many handbooks sport a big colour picture from esci on its cover 🙂

Joe Rodgers gives us a super-useful new resource. Chapter 1 (follow the steps in the figure caption) is a wonderful reflection on lessons from his lifetime of teaching statistics, then a brief summary of every chapter in the book.

At the publisher’s site for this book, click ‘preview this book’ (button may take a few moment to appear) under the pic of the cover. You can scroll through contents, etc, to the full text of Chapter 1.

The cover image is from Chapter 9, a great contribution by Chip Reichardt, a long-time friend at the University of Denver. I first called there for a quick visit back in 2009, when I was early in the writing of what became my first book, UTNS, just after my first YouTube video of the dance of the p values, and several years before Open Science burst onto the scene.

Chip was an enthusiastic user of the original ESCI that I was developing in MS Excel. (The version of ESCI used in UTNS is still available here.)

A Treasure Trove of Web Applets for Statistics Teaching

…that’s what Chip’s great Chapter 9 describes, together with his sage advice for selecting and using applets for a wide range of teaching aims. He gives this link to a list of all the applet links, so any of those more than two dozen applets is only a couple of clicks away.

esci-web Animating the Dance of the r Values

Below is a slightly edited version of what Chip writes about the dance of the r values.

Whatever you do, don’t overlook the applet at:

https://esci.thenewstatistics.com/esci-dance-r.html

(Hat tip to Gordon Moore, creator of esci-web, referred to here and whence came the book’s cover image.)

Pro tip: click the ‘?‘ top right in the control panel, it turns green, and then see pop-out tool tips as you hover the mouse over a control.

A large population of scores is presented as faint grey dots (just visible in the cover image) in a scatterplot. The applet then draws a random sample of data from the population and shows these as dark blue dots in the scatterplot (with any outliers in red), so you can see where the sample of data falls within the larger population of scores. You can click to add the regression line and r value for that sample.

Click the Take Sample button to see a different random sample from the population. Click Run Stop to see a sequence of random sets of blue dots dancing–the light grey dots of the population not moving–with the speed of dancing set by the slider. See also the regression line and the sample r value dance.

Use the sliders in the upper yellow and green panels to set sample size N and population correlation ρ (rho). Explore the effect of changing the N and ρ values

I hesitate to single out one applet from among many other excellent applets, but this applet is simply magnificent in showing how scatterplots, regression lines, and correlations vary across random samples of data. (Thank you Chip!)

The Animated r Heap

A few clicks and you can watch the sequence of sample r values drop down the screen and form the r heap, the empirical sampling distribution of r.

Control panel at left. In the bottom panel the checkboxes turn on display of the lower right panel (the upper scatterplot shrinks) and turn on display of the green dots for sample r values, which accumulate to form the sampling distribution of the r values. Red dots are r values whose 95% CIs (not shown) would not capture the population correlation value (ρ = .30, set by the slider in the upper green panel at left) and marked in the figure by the vertical blue line. Pop-out tool tips are on (the ‘?’ top right in the control panel is green): note the mouse pointer and tool tip in the figure.

Chip’s Chapter 9 is not included in the preview, but it’s well worth finding in your library–lots of great advice for great dynamic teaching demos.

Thank you Chip for this terrific new teaching resource.

Thank you Joe for highlighting esci on the cover!

Geoff

Psychologists Help Society in so Many Ways

Portrait of Ron Cumming, aeronautical engineer, psychologist, pioneer (with Ross Day) of human factors in Australia

Ten Years Applying Psychological Science Inside the U.K. Government is fascinating. It prompted this post and is at 6. below, but first a back story.

TL;DR. Main points:

  1. Ron Cumming, my father, pioneered human factors in Australia in the 1960s.
  2. Human factors, or user-centred design (UCD), is the design of devices and systems to be safe, easy, and effective for users.
  3. Ron and colleagues persuaded Victorian politicians to pass in 1970 the world’s first compulsory seat belt laws–road fatalities dropped immediately.
  4. Thaler & Sunstein’s Nudge gave examples of how simple changes in messages and systems can help people make better choices.
  5. I taught UCD at La Trobe for many years. Students had great fun identifying easy ways to improve the design of everyday things simply by watching and talking to real users. Some tell me they were inspired to go on and work in the field.
  6. In her fascinating column Ten Years Applying Psychological Science Inside the U.K. Government Carla Groom describes how nudge and UCD ideas can be used to make dramatic improvements to people’s lives. And how psychology PhDs can be effective leaders in such non-academic roles.

User Centred Design

A UCD classic, published in 1988, needed little updating for a new edition in 2012

A few examples in the first lecture and the students were saying “this is obvious, why isn’t everything done this way?” Don Norman‘s first few pages prompt the same reaction.

It should be obvious whether a door should be pushed or pulled, without labels. A first project: simply observe people as they approach different doors around campus.

Next, ask someone to open an unfamiliar microwave, or turn on the light on the left, or mute this mobile phone. Ask them to ‘think aloud’ as they approach the task, then watch without interrupting.

No RCT, control group, or fancy statistics required. Watch just a few users before–if you can–redesigning the object or system. Then do it again, and again. But you need real users–old, young, left-handed, neurodiverse, not speaking your language, perhaps disabled…

It can be frustrating. You quickly start noting poor design everywhere: why must I type my email in two different places? Why can’t I dislike all the options? Why do I need my glasses to figure out how to turn on the upper shower? …lots of scope for your psychology students to improve the world!

Landing Aircraft Safely

Ron Cumming, my father, was an aeronautical engineer researching crashes at the most dangerous moment in flying–landing. He decided it was largely a perceptual problem–the pilot perhaps needing to land with no visible horizon and only a single line of lights down one side of the runway.

Ron spent 1960 with human factors guru Paul Fitts at the University of Michigan studying psychology. Then he and psychologist Ross Day introduced human factors (ergonomics) to Australia.

Ron and colleagues developed T-VASIS in the 1960s. It was a simple array of lights each side of the runway that indicated to the pilot whether the aircraft was on the ideal glidepath, or needed to fly up or down a little. It was installed in many countries, and some airfields still offer it, despite the widespread use of modern radar systems.

T-VASIS: The view from the cockpit, coming in to land

Leading in Road Safety Was Just the Start

After leading the world with compulsory seat belts, in 1976 Victoria was also first with random breath tests of drivers, with .05 as the alcohol limit. Ron and Ross, as the two psychology professors at Monash University, set up the Monash University Accident Research Centre. The good work of MUARC continues.

In 1987 Victoria set up VicHealth as the world’s first health promotion foundation under legislation claimed to set the standard for international best practice by banning tobacco advertising and diverting those advertising dollars to fund anti-smoking campaigns and buy out tobacco sponsorship of sport and the arts.

The Slip-Slop-Slap campaign: Slip on a shirt, Slop on sunscreen, Slap on a hat,

In 1988 VicHealth funded launch of SunSmart, with the aim of changing behaviour to increase protection against UV in sunlight and thus reduce skin cancer. For example, with its Slip-Slop-Slap campaign.

These and other public behaviour change programs have saved numerous lives, averted much suffering, and reduced health costs. Psychologists continue to play prominent roles in all these programs.

Carla Groom Diagnosed as Autistic at 44

Dr Carla Groom, the then Head of Human Centred Design Science, U. K. Department of Work and Pensions was interviewed about her later-in-life formal diagnosis as autistic, at age 44. Again she is fascinating, here in recounting how the diagnosis explained for her so much about herself and how she worked.

She tells of the strategies she uses to support other neurodiverse people to be their most effective and, more generally, how she builds diverse–usually multi-disciplinary–teams, which tend to be better at problem solving.

Training PhDs to Be Effective in Non-Academic UCD and Nudge Work

Ten Years Applying Psychological Science Inside the U.K. Government offers lessons for how postgraduate education can be broadened to prepare PhDs to work and lead effectively beyond the academy. Including: study qualitative research methods, learn to write for a general audience as well as for academic journals, work in multi-disciplinary groups, and undertake messy real-world projects.

Such as helping pilots to land safely, or nudging everyone to wear a wide-brimmed hat and apply sunscreen.

Geoff

Brunch With Petra at the Museum

Petra Vaiglova is now Senior Lecturer in Archaeological Science at the Australian National University in Canberra. I posted here about meeting her for the first time, when she visited The University of Melbourne a few years back.

Geoff, Petra, and Stephen. Behind is a very early model Holden, manufactured in Australia, towing a caravan also of about 1950.

Since then she has published two especially notable open-access articles, the first being How can we improve statistical training in archaeological science?

Improving statistical training in archaeology science

Here’s the graphical abstract, the great work of Kathryn Killackey:

See the P.P.S. below for the full abstract.

Teeth, and ritual feasting a long long time ago

The second article, also open-access, is Transport of animals underpinned ritual feasting at the onset of the Neolithic in southwestern Asia, which reports a highly innovative study, led by Petra, of teeth from an Early Neolithic site in Iran.

I was recently in Canberra and, happily, could catch up with Petra and her partner, Stephen, for brunch at the National Museum cafe. I enjoyed a very good brunch, with animated and highly interesting discussion.

Geoff

P.S. Petra also very kindly made the trek from Canberra to join a rcent party marking my 80th birthday.

Yikes!

P.P.S. The abstract of Petra’s statistics article:

  • Raising the standard for statistical training in archaeology will improve the breadth and depth of archaeological science.
  • Improving statistical training can start by discussing five fundamental statistical concepts that archaeologists do not talk about enough.
  • Supervisors can help make statistical training more effective by advocating for statistical reform and Open Science.

The aim of this paper is to shine light on fundamental statistical concepts that archaeologists do not talk about enough. I argue that more deliberate discussion of these statistical ‘elephants in the room’ can have a positive impact on improving statistical training and on steering us away from perpetuation of poor research practices.

1) Statistical thinking should come first. This will help us break down some of the stigma around numbers and statistics, and set us up for building analytical frameworks that will provide the most informative answers to our research questions.

2) Descriptive and inferential statistics have different interpretative potential. This will clarify how we can move from using tools that only allow us to talk about our studied samples to using tools that enable us to draw inferences about the underlying populations from which the samples derived.

3) p values can be extremely variable. This will help spread awareness about the misuses and misconceptions of Null Hypothesis Significance Testing (NHST) and demonstrate the dangers of using significance thresholds to interpret data.

4) Statistical precision is not the same as measurement precision. This will bring attention to the many different types of uncertainties that are built into archaeological datasets (e.g., statistical precision, instrument measurement error, natural variation),.Recognising this is key for drawing reliable inferences from our data.

5) Meta-analyses and forest plots can be useful for synthesising previous research. This will help spread awareness about the benefit of meta-analyses for creating evidence-driven summaries of previous findings.

The discussion draws on examples from isotope archaeology, bioarchaeology, and organic residue analysis to illustrate how switching from a reliance on significance testing to a reliance on effect sizes can improve methodological rigour and the representativeness of our findings. The paper ends with a discussion of the roles and responsibilities of supervisors for creating an effective learning environment for statistical training. This includes, but is not limited to, acknowledging the problems of NHST and advocating for adherence to Open Science principles. Ultimately, the changes suggested in this paper will help us raise discipline-wide standards for quantitative training and improve both the breadth and the depth of archaeological research.

Simine Vazire Drops In for Dinner!

Of course I had to wear my APS uniform. You can just see, over Simine’s shoulder, the stunning artwork, portions of which appear on the cover of all my books.

It was great to see Simine last night, and to hear about the latest from APS and the world of journal editing. Her great love is still journal editing–very fortunately for psychology and indeed all of science. But when she wants quiet time for writing she takes to the road, finding quiet corners to use her laptop in coffee shops or, I was surprised to hear, breweries!

This last week she has slowly driven down from Sydney, coffee shop by coffee shop, to spend time with Fiona Fidler and with the MetaMelb group at The University of Melbourne.

Plastic plate, after 30 years of
trips through the dishwasher

She dropped in for a tuna pasta dinner, which grandkids Lucy and Zoe always request. Simine ate off this StatPlay classic picture of the green heap, long since renamed the mean heap.

These days you can explore the dances and many other goodies in esci web.

Enjoy!

The mean heap, from esci web

Geoff

P.S. On a personal note, we recently had a busy month to mark my 80th birthday (yikes!)–wonderful chamber music, two big parties, and now, to cap it off, a visit from Simine!

Research Priorities in Climate and Health Research

The best research method–of course–but also philosophy, ethics, global heating and health: What more could we want? This article, below, discusses all that and more. And, by the way, the lead author happens to be my son, Toby Cumming.

Here’s the abstract:
Toby Cumming

Rapid global warming is triggering a wide range of changes to the climate, and these changes are compromising many aspects of human health and well-being. As a research community, we lack the time and resources to investigate the efficacy of every possible climate adaptation strategy for protecting health. Thus, we require a logical and ethical framework to inform prioritization of our climate-health research efforts. In this paper, we propose a utilitarian approach: our research focus should be on adaptation strategies that provide the greatest health and well-being for the greatest number. The disability-adjusted life year (DALY) – equal to one year of healthy life lost – allows us to compare across markedly different health outcomes and adaptation approaches. Given the importance of cost-effectiveness in resource-constrained settings, we could prioritize adaptation approaches based on “cost per DALY averted”. Equally, we could base a priority ranking on “cost per quality-adjusted life year (QALY) gained”. A DALY- or QALY-based ranking would not be the end product, but a quantifiable first step to frame prioritization discussions. Adopting a utilitarian approach is useful in extending this frame beyond considering only the health of current-day humans to also consider the health of future humans and the suffering of non-human animals. While the approach does have limits – to ensure an equitable prioritization we need to consider aspects of fairness and justice, moral concepts that utilitarianism has some difficulty incorporating – we argue that it provides a helpful starting point in prioritizing the climate-health adaptation research agenda.

Finally

My congratulations to Toby and co-authors. For Toby, I happen to know this article draws heavily on his theoretical essay written, way back, as part of his Honours year in Psychology at La Trobe University. It was his outstanding performance in that year that took him to Cambridge for a PhD–and a wonderful three years of playing golf at the best courses up and down the British Isles. His PhD supervisor’s initial suggestion for his thesis title was “What I did when I wasn’t playing golf”. His recent passion project has been this book about the Australian golf courses designed in the 1950s and ’60s by Vern Morcom.

Whether you have a taste for climate change and health, or golf, please enjoy!

Geoff

Beyond the p Value: Reform Spreads Across the World and Across Disciplines

…as evidenced by this article from Brazil, which I’m delighted to see:

The article’s header

I salute Karen Grimmer, JECP co-editor, for publishing it, and for managing to make it Open Access. Karen happens to be a long-standing friend of mine who now continues the good work in ‘retirement’. I declare an interest: I was a referee for the ‘Beyond the p Value…’ article.

Note the innovative review, evaluation, and synthesis techniques developed by the authors. Here’s the Abstract:

Rationale

The p value has long been used as the primary criterion for statistical significance; however, its dichotomous interpretation has been increasingly criticized for oversimplifying uncertainty and distorting scientific inference, particularly in health and sports sciences.

Aims and Objectives

This study aimed to critically analyze the limitations of using the p value as the central criterion of statistical significance and to discuss more robust methodological alternatives for statistical inference.

Methods

A critical review was conducted using the PubMed/MEDLINE database covering the period from 2015 to 2025, complemented by citation tracking. Reviews, editorials, guidelines, and methodological essays that directly addressed the interpretation of p values and complementary metrics were included. A total of 46 articles were selected and evaluated using a self-developed critical appraisal checklist.

Results

Among the included studies, 38 (82.6%) explicitly criticized the isolated or dichotomous use of the p value, whereas eight adopted a more moderate position, supporting its use only when combined with confidence intervals, effect sizes, or Bayesian approaches. No article defended the p value as a standalone criterion for scientific decision-making. The most frequent recommendations involved abandoning the term “statistically significant,” prioritizing the estimation of effect magnitude and precision, and promoting the use of compatibility intervals, effect sizes, and Bayesian methods.

Conclusion

Overcoming the binary logic of p < 0.05 is essential to enhance transparency, reduce bias, and better align statistical practice with the scientific and clinical relevance of research findings, particularly in the health and sports sciences.

A Practical Summary

A particularly useful feature is a dot point summary of practical recommendations near the end:

  • Pose quantitative research questions (“to what extent…?”).
  • Report effect sizes with compatibility intervals as primary results.
  • Interpret uncertainty explicitly; avoid dichotomous terms.
  • Use estimation-focused graphics (e.g., Gardner–Altman, drapery plots, p value functions).
  • Preregister hypotheses and analysis plans.
  • Adopt cumulative reasoning (meta-analytic thinking).
  • Perform robustness and sensitivity analyses.
  • When applicable, evaluate hypotheses using an interval null.
  • Share data, materials, and code (Open Science practices).
  • For clinicians: emphasize magnitude and precision rather than thresholds.

Do pass this on to your clinical colleagues!

Geoff

Thankyou Jiangang! A Ten-Year Journey to Significance Roulette

Jiangang Xia is an enterprising professor at the University of Nebraska who, among many other things, teaches into China. He alerted me some years ago to the difficulty his students in China had accessing my videos because YouTube was blocked for them. So I mounted the videos at this OSF site.

As another enterprising step, Jiangang has recently been making a number of posts to LinkedIn. Below is one of these:

From a LinkedIn post by Jiangang Xia

Rethinking Quantitative Reasoning in Educational Research (4)

Jiangang Xia

Jiangang Xia

Associate Professor of Educational Administration at University of Nebraska–Lincoln

January 13, 2026

Geoff Cumming and the New Statistics: Estimation as a Way of Thinking

Some years only reveal their importance in hindsight. For me—and for statistics reform—2014 was one of those years.

That was the year I had just begun my academic career at the University of Nebraska–Lincoln. It was also the year Rex Kline visited UNL and delivered his keynote, “Hello, Statistics Reform.” (for Kline’s talk, please see my Post 2 for more details). At the time, I had no idea how much that visit would eventually shape my thinking, teaching, and research.

What I did not know then—and only came to appreciate much later—is that 2014 was also the year Geoff Cumming delivered his workshop “The New Statistics: Estimation and Research Integrity” at the APS Annual Convention in San Francisco.

I was completely unaware of that workshop.

In fact, I would remain largely unaware of Cumming’s work for several more years—even though it quietly passed through my academic life more than once.

Missed encounters (seen only in hindsight)

While preparing this post, I did something simple: I searched Geoff Cumming’s name in my old emails. That’s when I realized how often his work had crossed my path without fully registering.

In 2016, I co-chaired a dissertation in which the student cited Cumming’s influential article “The New Statistics: Why and How” (2014). At the time, neither of us fully grasped what the article was really asking us to reconsider. Although reform language appeared, the dissertation still relied on phrases like “marginally significant” for p values above .05—an indication that NHST logic remained firmly in place. Estimation had been encountered, but not yet learned as a way of thinking.

In 2017, I received an email from Routledge inviting faculty to request a desk copy of Introduction to the New Statistics. I didn’t request it. Another missed opportunity—one I only recognize now, looking backward.

These weren’t personal oversights so much as reflections of how deeply NHST was normalized in our training. Reform ideas were present, but the infrastructure for learning them—how to teach them and how to use them—was still thin.

From critique to alternative

It wasn’t until 2019, through Rex Kline’s work, that the larger picture finally came into focus for me. From there, I traced the reform movement backward—and Geoff Cumming’s role became unmistakable.

Cumming and his colleagues were not simply extending Jacob Cohen’s critique of null hypothesis significance testing. Cohen had already shown, powerfully, why NHST was flawed and had pointed toward alternatives. He also played a central role in the APA Task Force on Statistical Inference, whose report was released in 1999.

Unfortunately, Cohen passed away in 1998, before that work could be fully carried forward. When the APA’s 5th edition was published in 2001, many of the reform ideas were only partially adopted, and everyday research norms remained largely unchanged.

What Cumming did next was different.

Through decades of scholarship—culminating in The New Statistics—he articulated a coherent alternative centered on estimation rather than binary decisions, emphasizing effect sizes, confidence intervals, precision, uncertainty, and cumulative evidence. By the New Statistics, Cumming meant an estimation-centered approach to inference—focusing on effect sizes, confidence intervals, and uncertainty, and on combining evidence across studies through meta-analysis—rather than making binary decisions based on p values alone.

From alternative to institutionalization and teaching

That work did not remain theoretical.

In 2008, Cumming was invited to join the small working group responsible for statistical reporting standards in the APA’s 6th edition, where he was a driving force behind the requirement that effect sizes and confidence intervals be reported for every research question—helping move estimation from an optional supplement to a core reporting standard.

Just as importantly, Cumming took the initiative to teach this alternative—by writing new articles (2014), new textbooks (2013, 2017, 2024), offering workshops (2014 APS), and developing demonstrations (ESCI) aimed at broad audiences, not just methodologists.

Why experience matters in teaching reform

Along the way, Cumming recognized something more fundamental: logical arguments alone were not enough.

As he later reflected, perhaps NHST had become “the researcher’s heroin—an addiction impervious to reason.”

If that was true, persuasion would require more than explanation. It would require experience.

This insight shaped his teaching. Rather than debating p values in the abstract, Cumming showed researchers—often viscerally—how unstable they are through demonstrations such as the dance of the p values, p intervals, and significance roulette. The goal was not just to convince the mind, but to engage the gut—to help researchers feel uncertainty rather than deny it.

A decade later, the field itself began to catch up. In 2024, Cumming’s 2014 article “The New Statistics: Why and How” received SAGE’s 10-Year Impact Award, recognizing research whose influence endures well beyond the standard citation window. The award was a reminder that reform ideas are often understood slowly—resisted early, adopted unevenly, and acknowledged only after they have quietly reshaped teaching and research norms.

Full circle: from missed encounters to transformed practices

What changed for me after 2019 was not just what I read—but how I taught, mentored, and conducted research.

In 2022, I shared Cumming’s 2014 article with my doctoral advisee Amanda. Her dissertation became the first I supervised to fully abandon significance language, adopting estimation-based interpretation throughout. That same year, a manuscript my student Cailen and I submitted—published in 2023—explicitly drew on Cohen, Kline, Cumming, and the ASA (2016) statement. It was my first research article grounded fully in estimation thinking.

And in a quiet but meaningful full circle, when I needed materials for my 2024 summer teaching in China, Geoff Cumming himself shared his 2014 APS workshop videos with me—materials I have since used both internationally and in my quantitative methods courses at UNL.

Looking back now, the reform was happening all around me in 2014.

I just wasn’t ready to see it.

Perhaps that is how intellectual change often works—not as a single revelation, but as a series of missed encounters that eventually align. And perhaps genuine reform requires not only better arguments, but better ways of helping researchers experience uncertainty.

For readers who want to explore further

(All materials shared with permission; enormous credit to Geoff Cumming, Bradley Dean, and Robert Calin-Jageman.)

Dear colleagues,

  • When did you first encounter Geoff Cumming’s work—if at all?
  • Have you seen estimation treated as an add-on, rather than a way of thinking?
  • What important ideas did you meet early, but only understand much later?

#Cumming #NewStatistics #EstimationThinking #QuantitativeMethods #EducationalResearch

Thank you Jiangang!

Geoff

A Statistics Textbook for the AI Era

Miodrag Lovrić is an enormously energetic statistician and educator. He persuaded 700 scholars from 110 countries to contribute to the massive four-volume second edition of the International Encyclopedia of Statistical Science (Springer, 2025).

Now he is close to completing Statistical Thinking for the AI Era, which comprises Vol. 1: Foundations and Inference, and Vol. 2: Advanced Methods–each with about 15 chapters and 520 pages. The chapters I have seen indicate it’s a terrific textbook.

Miodrag has lived and worked on at least four continents and takes a wholeheartedly global focus in both those major works. The textbook starts with Forewords from many continents. Miodrag generously invited me to contribute a foreword from Australia–which appears below. As you see, I slipped in mention that I’m an Antarctic tragic, even if it’s a big claim that I speak also from that seventh continent!

Foreword From Australia and Antarctica

I’ve been an Antarctic tragic since I was a boy. I’ve participated in citizen science during two voyages south. I’ve observed glacial retreat and penguin colonies dying. I read about giant datasets analysed by AI-assisted statistical modelling of changes to sea ice, weather, and much else including terrifying tipping points that will determine catastrophic changes our children must endure.

I’ve given a well-received statistics talk to researchers at the Institute for Marine and Antarctic Studies in a building on the wharf where Antarctic supply ships berth. All around us were whale skulls, ancient sledges, and other memorabilia. I feel I can claim at least a little justification for offering this piece on behalf of the seventh continent, as well as speaking from Australia.

Statistics is about communication and therefore the province of psychologists. Researchers publish representations of their results and readers, ranging from researchers to politicians, policymakers, and ordinary people, draw conclusions from those representations. However, vast industries misrepresent research results as they seek to persuade us to support autocrats, eat unhealthy food, buy things we don’t want that trash the planet, and burn fossil fuels to destroy the chance our children will inherit a liveable world.

My students and I studied people’s misconceptions of basic statistical concepts. That was statistical cognition, which has now developed hugely into metascience—a wonderfully interdisciplinary field that investigates how research is done, and should be done to be more trustworthy. For example, it studies how paper mills use AI tools to generate fake manuscripts that desperate researchers buy then submit to journals. Quality journals use AI tools and much editorial expertise to try to weed out fakes before these criminally pollute the research literature—much of it on life-and-death issues.

Around 2014 Open Science emerged: replication is central; this requires meta-analysis to synthesise results; this requires point and interval estimates, and statistical significance is irrelevant and often damaging. My statistics textbooks advocate estimation and meta-analysis. The intro book, with Bob Calin-Jageman, integrates estimation and Open Science all through. The second edition (Routledge, 2024) includes Bob’s wonderful open source software that’s ideal for beginners and researchers aiming for good estimation and Open Science practices. At thenewstatistics.com is more information, also the significance roulette simulation, which dramatizes a highly misleading feature of p values that few researchers appreciate. The good news is that these new approaches are found accessible by students and are a delight to teach.

I conclude that everyone should have a basic understanding of evidence and statistics. This book is broad in scope, well-informed, and future-focussed. For example, it explains traditional, Bayesian, and randomisation frameworks for estimation, any of which can support Open Science. As I read, I’m often applauding the approaches Professor Lovrić has chosen. With this book he is making an enormous contribution to statistics education around the world.

Geoff

The International Research Integrity Conference

Simon Gandevia, a giant in the research integrity community, assembled an international A-list of incredible speakers–plus me–for this Conference (IRIC) recently in Sydney.

Diversity of disciplines, but a common focus

There were more than 100 registrants, so the presentation room was usually packed. A striking feature was our enormous diversity of disciplines: medicine, law, neuroscience, statistics, psychology, journalism, education… and more. Our common focus was the integrity of research, Open Science, and meta-research (research on how research is done). Most presentations to the conference were either an alarming assessment of how the research literature has been polluted, or a (sometimes) encouraging evaluation of efforts to improve integrity.

We felt strongly about so many things, with emotions swinging wildly. Mainly we felt enormous pride in group cohesion across such a diversity of backgrounds, and conviction that our mission–the very trustworthiness of research–is enormously important.

Paper mills, fakery, and more

But then we’d despair when a speaker documented the extend of deadly pollution spreading in the scientific literature, from old-fashioned fakery to plagiarism and fake manuscripts sold by paper mills. Anyone wanting to bulk up their cv can buy an AI-generated manuscript and submit it to a dozen journals. The best journals will identify the fakery, these days using AI to help, but some ‘vanity’ journal will publish the manuscript after a bogus peer review process–of course for a hefty payment. What criminality! What waste, especially of the time and effort of expert editors!

The memorable conference dinner was held on the roof deck of the stunning new Museum of Contemporary Art. Champagne in hand, sun setting on the harbour bridge to the left, opera house to the right: what’s not to like?

Significance Roulette released to the world

I’m delighted to say my brief presentation was very well received. Most importantly, I released to the world for the first time Bradley Dean‘s Significance Roulette that you can download, save, and run in your browser. Or simply go to esci web and select the significance roulette menu item to see and run the roulette wheel.

From here you can download my powerpoint slides, also a pdf summary, and the .html file you can save to your computer and run in your browser to enjoy significance roulette.

Read more about the rationale for significance roulette here.

Good luck at the p value casino!

Geoff

Running Away From Significance As Hard As It Can?

How many times have you read that p = .062 is “approaching” significance? How do you know it’s not sprinting as hard as it can away from statistical significance? (Not my idea, unfortch, but a quote from some insightful person way back.)

I’ve just rediscovered the wonderful list of something like 500 ways “approaching” has (allegedly!) been expressed.

Weep!

Geoff