Brunch With Petra at the Museum

Petra Vaiglova is now Senior Lecturer in Archaeological Science at the Australian National University in Canberra. I posted here about meeting her for the first time, when she visited The University of Melbourne a few years back.

Geoff, Petra, and Stephen. Behind is a very early model Holden, manufactured in Australia, towing a caravan also of about 1950.

Since then she has published two especially notable open-access articles, the first being How can we improve statistical training in archaeological science?

Improving statistical training in archaeology science

Here’s the graphical abstract, the great work of Kathryn Killackey:

See the P.P.S. below for the full abstract.

Teeth, and ritual feasting a long long time ago

The second article, also open-access, is Transport of animals underpinned ritual feasting at the onset of the Neolithic in southwestern Asia, which reports a highly innovative study, led by Petra, of teeth from an Early Neolithic site in Iran.

I was recently in Canberra and, happily, could catch up with Petra and her partner, Stephen, for brunch at the National Museum cafe. I enjoyed a very good brunch, with animated and highly interesting discussion.

Geoff

P.S. Petra also very kindly made the trek from Canberra to join a rcent party marking my 80th birthday.

Yikes!

P.P.S. The abstract of Petra’s statistics article:

  • Raising the standard for statistical training in archaeology will improve the breadth and depth of archaeological science.
  • Improving statistical training can start by discussing five fundamental statistical concepts that archaeologists do not talk about enough.
  • Supervisors can help make statistical training more effective by advocating for statistical reform and Open Science.

The aim of this paper is to shine light on fundamental statistical concepts that archaeologists do not talk about enough. I argue that more deliberate discussion of these statistical ‘elephants in the room’ can have a positive impact on improving statistical training and on steering us away from perpetuation of poor research practices.

1) Statistical thinking should come first. This will help us break down some of the stigma around numbers and statistics, and set us up for building analytical frameworks that will provide the most informative answers to our research questions.

2) Descriptive and inferential statistics have different interpretative potential. This will clarify how we can move from using tools that only allow us to talk about our studied samples to using tools that enable us to draw inferences about the underlying populations from which the samples derived.

3) p values can be extremely variable. This will help spread awareness about the misuses and misconceptions of Null Hypothesis Significance Testing (NHST) and demonstrate the dangers of using significance thresholds to interpret data.

4) Statistical precision is not the same as measurement precision. This will bring attention to the many different types of uncertainties that are built into archaeological datasets (e.g., statistical precision, instrument measurement error, natural variation),.Recognising this is key for drawing reliable inferences from our data.

5) Meta-analyses and forest plots can be useful for synthesising previous research. This will help spread awareness about the benefit of meta-analyses for creating evidence-driven summaries of previous findings.

The discussion draws on examples from isotope archaeologybioarchaeology, and organic residue analysis to illustrate how switching from a reliance on significance testing to a reliance on effect sizes can improve methodological rigour and the representativeness of our findings. The paper ends with a discussion of the roles and responsibilities of supervisors for creating an effective learning environment for statistical training. This includes, but is not limited to, acknowledging the problems of NHST and advocating for adherence to Open Science principles. Ultimately, the changes suggested in this paper will help us raise discipline-wide standards for quantitative training and improve both the breadth and the depth of archaeological research.

Beyond the p Value: Reform Spreads Across the World and Across Disciplines

…as evidenced by this article from Brazil, which I’m delighted to see:

The article’s header

I salute Karen Grimmer, JECP co-editor, for publishing it, and for managing to make it Open Access. Karen happens to be a long-standing friend of mine who now continues the good work in ‘retirement’. I declare an interest: I was a referee for the ‘Beyond the p Value…’ article.

Note the innovative review, evaluation, and synthesis techniques developed by the authors. Here’s the Abstract:

Rationale

The p value has long been used as the primary criterion for statistical significance; however, its dichotomous interpretation has been increasingly criticized for oversimplifying uncertainty and distorting scientific inference, particularly in health and sports sciences.

Aims and Objectives

This study aimed to critically analyze the limitations of using the p value as the central criterion of statistical significance and to discuss more robust methodological alternatives for statistical inference.

Methods

A critical review was conducted using the PubMed/MEDLINE database covering the period from 2015 to 2025, complemented by citation tracking. Reviews, editorials, guidelines, and methodological essays that directly addressed the interpretation of p values and complementary metrics were included. A total of 46 articles were selected and evaluated using a self-developed critical appraisal checklist.

Results

Among the included studies, 38 (82.6%) explicitly criticized the isolated or dichotomous use of the p value, whereas eight adopted a more moderate position, supporting its use only when combined with confidence intervals, effect sizes, or Bayesian approaches. No article defended the p value as a standalone criterion for scientific decision-making. The most frequent recommendations involved abandoning the term “statistically significant,” prioritizing the estimation of effect magnitude and precision, and promoting the use of compatibility intervals, effect sizes, and Bayesian methods.

Conclusion

Overcoming the binary logic of p < 0.05 is essential to enhance transparency, reduce bias, and better align statistical practice with the scientific and clinical relevance of research findings, particularly in the health and sports sciences.

A Practical Summary

A particularly useful feature is a dot point summary of practical recommendations near the end:

  • Pose quantitative research questions (“to what extent…?”).
  • Report effect sizes with compatibility intervals as primary results.
  • Interpret uncertainty explicitly; avoid dichotomous terms.
  • Use estimation-focused graphics (e.g., Gardner–Altman, drapery plots, p value functions).
  • Preregister hypotheses and analysis plans.
  • Adopt cumulative reasoning (meta-analytic thinking).
  • Perform robustness and sensitivity analyses.
  • When applicable, evaluate hypotheses using an interval null.
  • Share data, materials, and code (Open Science practices).
  • For clinicians: emphasize magnitude and precision rather than thresholds.

Do pass this on to your clinical colleagues!

Geoff