Open Science Practices: Patchy Progress in Two Psychology Journals

(Revised 18 March 2022, to add comments about replication, and the effectiveness of journal-specific guidelines, and badges.)

What Progress With Open Science? In Brief:

Judging from Psychological Science (PS) and the Journal of Experimental Psychology: General (JEPG), during 2013-2020:

  • Reporting of Confidence Intervals (CIs) increased markedly from 2013 to 2015 🙂 but has plateaued at around 60% of articles since then 🙁
  • CIs are still rarely explicitly used to inform interpretation of results 🙁 🙁
  • Use of Effect Sizes (ESs) to inform interpretation has increased steadily to around 50% 🙂
  • Provision of Open Data and Analysis Code, and Open Materials has increased steadily, from near zero to around 50% 🙂
  • Use of Preregistration has increased from near zero to around 30% 🙂
  • Use of NHST remains almost universal 🙁
  • Overall, only 1.9% of articles reported replications 🙁 and there were no registered reports.
  • Larger changes in PS than JEPG suggest strong journal-specific policies and offering Open Science badges can be effective 🙂 🙂

Conclusion: Open Science has made enormous strides since 2013, but there is still a long way to go. Keep at it!

Our Two Studies

In 2017, David Giofrè and colleagues reported a study of the frequency of use of various statistical and Open Science (OS) techniques from the start of 2013 to the end of 2015, in PS and JEPG. We recently uploaded a preprint that updates the picture, for articles published in those two journals from the start of 2016 to the end of 2020.

From 2013 to 2015

In January 2014, Psychological Science famously announced dramatic changes to its instructions to authors. Erich Eich, then Editor-in-Chief, explained in this editorial. Among other changes, use of The New Statistics was strongly encouraged, reliance on NHST was discouraged, fully detailed reporting was required, and badges could be earned by providing open data, or open materials, or by reporting preregistered research. The figure below shows the proportions of articles using various practices each year from 2013, before the changes, to 2015, when there had been time for the changes to influence what was published.

We included JEPG for comparison. Its instructions to authors were not as detailed, and relied largely on a general reference to the APA Publication Manual. There were no marked changes to its requirements during the period.

Proportions of articles in the two journals that used various statistical and Open Science practices, 2013-2015,

The good news includes: In PS, use of CIs (item 2 in the figure) and provision of open data (8) and open materials (9) increased dramatically from 2013 to 2015. Note that OS badges to acknowledge 8 and 9 (also 10) were introduced in 2014. Some justification for chosen sample size (6) and explanation of any data exclusions (7) also increased strongly. There were similar but generally not so marked changes in JEPG, consistent with PS’s strong journal-specific guidelines and offering of badges. JEPG changes perhaps reflected the rapid spread of OS consciousness occurring back then, even without specific changes to that journal’s policies.

The bad news includes: NHST (1) remained close to universal; use of CIs for interpretation and discussion of results (4) remained very low; and preregistration (10) was very rarely used.

From 2016 to 2020

We used the same procedure to assess practices during 2016-2020. We added one practice: the provision of data analysis code. This table reports percentages of empirical articles that used the various practices, for each journal in each year:

Overall, CI use (row 2) held up but hardly increased further, NHST (1) remained near-universal (boo!), and use of CIs for interpretation (4) remained rare (double boo). It’s great to see that several desirable practices all steadily increased, including: use of ESs for interpretation (5); provision of open data (8), open materials (9), and open code (11); and preregistration (10).

Explaining choice of sample size (6) and data exclusions (7) generally increased. Row 3 refers to internal meta-analysis, meaning meta-analysis of two or more studies reported in the article itself. It was almost unknown in 2013 and more recently occurred in as many as around 10% of articles–that’s a great development.

Reporting a replication of a previously published study wasn’t one of the practices we investigated in detail, but we can report that, overall, only 1.9% of articles reported such a replication, and none was a registered report.

As in the earlier period, PS generally did better or much better than JEPG, consistent with journal-specific guidelines and badges being effective ways to bring about OS improvements.

Looking back to the picture in 2013, it’s clear that, at least judging from these two leading journals, we’ve come a very long way towards more open and trustworthy published research. But there’s still much progress to be made, especially by encouraging replication, and explicit use of ESs and, in particular, CIs to inform interpretation of results. Yep, the new statistics simply make more sense–as intro students keep telling us.

Geoff

Giofrè, D., Cumming, G., Fresc, L., Boedker, I., & Tressoldi, P. (2017). The influence of journal submission guidelines on authors’ reporting of statistics and use of open research practices. PLOS ONE. https://doi.org/10.1371/journal.pone.0175583

Giofrè, D., Boedker, I., Cumming, G., Rivella, C., & Tressoldi, P. (2022, submitted for publication). The influence of journal submission guidelines on authors’ reporting of statistics and use of open research practices: Five years later. https://osf.io/preprints/metaarxiv/8ya3m

Hallelujah! Physiotherapists Join the Christmas Choir of Estimation Angels!

I woke on Christmas morning to a message from Bob with this link, to this editorial:

Note the number of Physio Journals represented in the authorship!

Not just a single journal, but journal editors for a whole discipline! Here’s the first para, highlights added:

Null hypothesis statistical tests are often conducted in healthcare research, including in the physiotherapy field. Despite their widespread use, null hypothesis statistical tests have important limitations. This co-published editorial explains statistical inference using null hypothesis statistical tests and the problems inherent to this approach; examines an alternative approach for statistical inference (known as estimation); and encourages readers of physiotherapy research to become familiar with estimation methods and how the results are interpreted. It also advises researchers that some physiotherapy journals that are members of the International Society of Physiotherapy Journal Editors (ISPJE) will be expecting manuscripts to use estimation methods instead of null hypothesis statistical tests.

instead of…” — that’s music to my ears! What a wonderful Christmas gift!

This is from the text, highlights added (p. 3):

ISPJE member journals’policy regarding the estimation approach

The executive of the ISPJE strongly recommends that member journals seek to foster use of the estimation approach in the papers they publish. In line with that recommendation, the editors who have co-authored this editorial advise researchers that their journals will expect manuscripts to use estimation methods instead of null hypothesis statistical tests.

Support, Resources, References

The editorial includes brief explanations, useful summary tables, and 38 highly useful references, including 4 of Bob’s and/or mine.

I am so, so happy to see physiotherapy join the choir!

A very happy estimation Christmas to all,

Geoff

MRI Workshop Videos, Including Two Short ‘p Values Suck’ Talks

I recently posted about MRI Together. It was a great global zoomfest, and now videos of the talks are online.

The Videos

The YouTube site with all the videos is here, but it’s easiest to scan the program and click on any talk title to go to the video.

It’s clear that many in the MRI community have been working on Open Science issues for a while. It was great to see lots of talks about software for statistical analysis of scans, and the challenges of making analyses reproducible–and of finding good ways to make code and the highly complex data sets open, and readily usable by others.

I’ll mention just a few videos below.

An Intro

Cassandra Gould van Praag gives an engaging intro talk here. She encountered a range of OS issues when her PhD examiners required modifications to her thesis. She bravely tells the story of the work and revisions she had to undertake. She must have taken the lessons to heart: Her current job is to develop and promote the open science infrastructure of Oxford Neuroscience.

Two ‘p Values Suck’ Talks

Valentin Amrhein gave a rousing talk titled ‘P Values and the Replicability of Results‘. The video is here. He has some well-chosen graphics and some striking graphs, for example of the wide range of likely p values in various situations. You won’t be surprised to hear that I heartily agree with almost everything he said.

My talk was titled The New Statistics for Reproducible Science The video is here. My three take-home messages were:

  1. p values are highly unreliable, never trust them
  2. Adopt Open Science practices, planned analyses
  3. esci, Bob’s new software for estimation, meta-analysis, teaching, and more, is available now (in beta), and gaining thousands of installs each month. It includes great graphs with confidence intervals. It’s a pleasure to teach with it.

Brian Nosek Rounds It Off

Talk title: Publishing and Sharing Open Science. The video is here. Brian does a typically neat and persuasive job, of course with lots of evidence.

Lots more gems to discover: Just browse the program.

Geoff

MRI Analysis, Now With Open Science

It’s fabulous to see yet one more research field, MRI and fMRI, jumping on board with Open Science. The Workshop runs 13-17 December, 2021, and the site is here.

Recall the dead salmon? (tiny.cc/deadsalmon, ITNS p. 485.) Back then, in 2009, a common way to analyse fMRI data relied on p values for each of many thousands of voxels. Apply this analysis to a dead salmon shown two different types of pictures and find part of its (totally dead) nervous system lit up on the analysis screen! It was a massive Type I error, of course, caused by inadequate correction of p values given by the gazillion voxel comparisons. Make a more appropriate correction and all we see is noise. The study deserved its 2012 IgNobel Prize.

Analysis of MRI data has come a long way since 2009, partly prompted by the dead salmon. More recent discussions, for example here and here, include consideration of the basic Open Science issues, including p-hacking, cherry-picking, unplanned analyses, and lack of replication.

The Timetable (Program)

It’s here. Near the top, select your timezone. There are sessions around the 24 hours, grouped to fit the waking hours in Atlantic, Pacific, Indian, and Caribbean zones. It looks to me like a wonderfully broad take on Open Science and Reproducibility. I recognised only a few of the presenters, including:

Valentin Amrhein: P-values and the replicability of results 

Brian Nosek: Publishing and Sharing Open Science 

I’m giving a brief (20 min, including Q+A) talk: The new statistics for reproducible science  in a session titled Study design and interpretation. The reproducibility crisis, running 11.00 to 13.00 on 15 Dec, those times being UTC+11, which includes Sydney and Melbourne. I’ll post my slides in due course.

Geoff

One More Time: p Values are Scarily Unreliable

Replicate a study and you are highly likely to get a very different p value. Scarily different. The sampling variability of p values is so great that no p value deserves our trust.

Yes, I know, that’s been my mantra for more than a decade, but Steve Lindsay has just given me an excuse to beat the drum again. This is his great article:

It’s clear that Lindsay (2020) (preprint here) has been written by a practical, practising researcher. For example, there’s advice about references to help any new arrival in your lab get up to Open Science speed, and discussion about developing a laboratory manual to help everyone adopt good, systematic OS practices. His seven steps are an operationalisation of many such practices.

Possibly broader and more detailed accounts of what OS needs were given by Asendorpf et al. (2013), and Munafò et al. (2017). Dorothy Bishop (2020) gave a particularly compelling account, with a focus on overcoming the cognitive biases that make adopting OS practices a challenge.

Dance of the p Values

To illustrate his discussion of variability of ESs over replications, Steve Lindsay uses a figure (below) of the dance of the CIs and dance of the p values.

Screenshot from ESCI for UTNS. The blue and red distributions depict control and experimental populations, whose means differ by Cohen’s δ = 0.5 (or 10 raw points). The 20 open blue and 20 open red circles were randomly sampled from those populations. Each solid green circle represents the result of a simulated experiment: A raw effect size (experimental mean minus control mean) from random draws from the two populations, with the 95% CI around that ES estimate. The topmost (most recent) solid green circle represents the difference between the solid blue and red circles, the group means. In that random draw, the difference between conditions was not statistically significant (p = .229, as shown at left, corresponding to that CI considerably overlapping the vertical H0 line), a type II error. Of the 25 simulated experiments shown here, three came out in the wrong direction because of random sampling error; seven experiments detected the effect (p < .05, CI fully to the right of the H0 line) and every one of those overestimated the size of the effect.

CIs Make Uncertainty Salient

In the figure it’s easy, with a bit of practice, to focus on a single CI, note where it falls in relation to the vertical line (H0 of zero difference), and eyeball the p value. (ITNS Chap 6 has handy guidelines.) The two dances are telling us the same story.

But, and it’s a huge ‘but’, any single CI—which is all we know in real research life—gives us information about how wide, how frenetic the dance is. It makes the uncertainty salient. In stark contrast, any single p value tells us virtually nothing about the dancing p values. Its single value, perhaps reported to three decimal places, gives a seductive, but illusory, sense of certainty, even though it could have been very different. Yep, p is not to be trusted!

Software for the Dance of the p Values

I used ESCI in Excel 2003 (the best version ever, imho!) for the original dance of the p values video, also for this later video. It ran beautifully quickly. Unfortunately, Excel 2007 runs very much more slowly, for example in ESCI for UTNS, 2007 version, the replications plod down the screen.

Happily, we now have esci web, thanks to Gordon Moore. The first component is dances. Click ‘?’, top right in left panel, to turn on tips. Then have a play. For the dance of the p values, click red ‘9’, bottom sub-panel. Turn on sound and use the slider to adjust the speed, which can be too fast. Enjoy!

p Values Unreliable in Virtually Any Situation

Have I chosen sample sizes and population ES, δ, to give a particularly dramatic dance of the p values? You can use ESCI for UTNS, or esci web, to investigate. You will discover, for example, that larger sample sizes and/or larger δ, give average p that’s smaller (higher power, so more replications give p < .05), but there is still crazy dancing of the values of p.

For virtually any situation, the replication to replication variation in p is scarily large. Whenever you see a p value reported, remind yourself that the study could easily have given a quite different value, and that an exact replication is likewise likely to give a very different value. Any p value should be regarded as very fuzzy indeed.

What if We Don’t Know the Population ES, But Only the p Value?

The dance of the p values assumes that we know the population ES, δ, and N. However, in real research life we never know δ. In Cumming (2008) I developed the formulas for the sampling distribution of the p value for two situations:

  1. δ and sample size(s) are known, as in the dances of the p values above; and
  2. all we know is the p value given by an initial study.

For 2, we assume a single replication, everything exactly the same as in the initial study except that a new sample is taken. (Also assumes sample sizes are not small.) I refer to the p value given by such a replication as replication p.

The article includes figures showing the two distributions for a range of cases. For the second case, the figure shows the distribution of replication p following a specified p value found in the initial study.

To my surprise, I found that it’s perfectly possible to find and picture the distribution of replication p: Tell me only p from your study, and I’ll tell you the chance that an exact replication will give, for example, p < .05. Whatever the power, the N (if not small), and the true ES.

Not surprisingly, the distribution of replication p is very wide. Yes, following an initial p = .01 you are likely, on average, to get smaller replication p than following initial p = .05. But in both cases there’s vast uncertaintly—almost any replication p may occur.

Significance Roulette

How could I dramatize the almost unbelievably large amount of uncertainty in replication p? A mere diagram of the sampling distribution lacks punch. I divided the area under the distribution curve into 38 equal areas, because that’s the number of slots on one common roulette wheel. I used the p value at the centre of each area to represent that area. I arranged those 38 p values haphazardly around a wheel, in fact a roulette wheel. This figure is the wheel for initial p = .05:

After an initial study gives p = .05, a replication, just the same but with a new sample, will give replication p as shown by the wheel. The distribution of replication p is illustrated at left in terms of five conventional intervals. All p values are two-tailed.

If you find an initial p = .05, then to find what p an exact replication is likely to give, spin the wheel. Each of the 38 values is equally likely. At left is a summary, in terms of five conventional intervals of p values. You have a 7/38 chance of *** (p < .001, bright red circles) and 15/38 chance of p > .10 (deep blue circles). A mere tiny flip of the wheel and, instead of p = .39, as pictured, you might have obtained .02, or maybe ** or ***.

Here’s the wheel for initial p = .01:

As for the previous figure, but now for initial p = .01.

As we expect, replication p values are on average smaller than for the first wheel, but there’s still enormous variation. Chance of *** is now 12/38 and chance of p > .10 is now 10/38. Once again there’s vast uncertainty ☹

Videos for Significance Roulette

For better explanation than my brief sketch above, see two videos, here (tiny.cc/SigRoulette1) and here (tiny.cc/SigRoulette2). Incidentally, the wheel runs in Excel 2003 but, despite great effort, I haven’t been able to get it running in later Excel. If anyone would like to build it in some more modern language, preferably to run on the web, that would be great. Please let me know.

Beyond p Values

I’ve long argued, for example in Cumming (2014), that p values are rarely needed and that almost always we’re better off if we simply don’t report them, and don’t see them in published articles. They should simply be left to whither away. The only reason students should need to learn about them is to make sense of old research literature.

However, NHST and p values seem to be the researcher’s heroin. For most of us, and the published literature, they are deeply embedded and it seems very challenging to overcome the addiction. Rational argument has not proved effective. Can dramatization of unreliability do better?

“To What Extent…?”

For years I have been schooling myself never to ask “I wonder whether…” even when daydreaming. It always has to be “I wonder to what extent…”. Give it a try. That’s the step from dichotomous to estimation thinking. Our world simply isn’t black-white, but a gazzillion shades of grey, not to mention colours.

Meanwhile, enjoy seeing the wheel spin, and musing about what it tells us.

May all your confidence intervals be short!

Geoff

Asendorpf, J. B., et al. (2013). Recommendations for increasing replicability in psychology. European Journal of Personality, 27(2), 108-119. https://doi.org/10.1002/per.1919

Bishop, D. V. M. (2020). The psychology of experimental psychologists: Overcoming cognitive constraints to improve research: The 47th Sir Frederic Bartlett Lecture. Quarterly Journal of Experimental Psychology, 73(1) 1–19. https://doi.org/10.1177/1747021819886519

Cumming, G. (2008). Replication and p Intervals: p values predict the future only vaguely, but confidence intervals do much better. Perspectives on Psychological Science, 3(4), 286–300. https://doi.org/10.1111/j.1745-6924.2008.00079.x

Cumming, G. (2014). The New Statistics: Why and how. Psychological Science, 25(1), 7-29. https://doi.org/10.1177/0956797613504966

Lindsay, D. S. (2020). Seven steps toward transparency and replicability in psychological science. Canadian Psychology/Psychologie canadienne, 61(4), 310–317. https://doi.org/10.1037/cap0000222

Munafò, M. R., et al. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, 0021. https://doi.org/10.1038/s41562-016-0021

Open Science: A Detailed Telling of the Story So Far

Concern about publication bias apparently goes back at least as far at Robert Boyle in the 17th century. That’s Barker Bausell’s starting point for his highly detailed account of the replication crisis and rise of Open Science. Bausell, the author of a number of books, brings a broad perspective, across many disciplines beyond only psychology, to the story he tells so engagingly. Frequent mentions of health and medical research reflect one aspect of his expertise.

The book is The Problem With Science: The Reproducibility Crisis and What to Do About It. You can read the first couple of dozen pages here, or go here and click at left for a brief summary of each chapter.

Chapter 2 is appropriately critical in discussing NHST and p values, and false positives. Recommendations focus on using larger and higher-power studies, ESs, and α values much smaller than .05. I, of course, wish for recommendations against using p values at all, and in favour of estimation—for single studies and as the basis for meta-analysis. That’s surely at the core of Open Science practice.

Chapter 3 discusses a monumental list of 23 (!) QRPs—and some of those actually comprise more than one practice.

All through there’s much that will be familiar to many readers of this blog. We come across HARKing, forking paths, and Bem, Bargh, and Cuddy. But we also hear about cold fusion, and other goodies (?) from chemistry, physics, and more. I love the whole-of-science perspective. Of course, there’s lots of discussion of replication. Also of publishing practices, and the terrible contingencies that researchers have traditionally faced.

For any particular problem there’s discussion of the full sequence: the problem’s origins, investigations of how serious and widespread it is, what’s needed to solve it, then finally how we’re going so far with implementing the necessary changes. This sequence reminds me of Replicability, Robustness, and Reproducibility in Psychological Science, by Brian Nosek et al., which is a detailed discussion of exactly what the title states and developments over the last decade or so. This is in press for The Annual Review of Psychology and the final version is available here as a preprint. By comparison, Bausell ranges over many disciplines.

Education

Bausell’s brief final chapter looks forward. I like his emphasis on the incorporation of Open Science practices into education—I would say that, wouldn’t I, with ITNS2 designed to be a resource for just that. He takes an optimistic position, judging that, given sufficient support by researchers, positive developments in research practices are likely to flourish, even if success is not (yet) assured. Good to see; I dearly hope he’s correct.

While discussing the challenge of overcoming ingrained beliefs that ‘statistically significant’ is equivalent to ‘notable’, Bausell writes “the best educational option is probably to keep the sanctity of obtaining p-values at all costs from being learned in the first place” (p. 264). I couldn’t agree more! Again, I’d go further and advocate starting with estimation, then mentioning p values, if at all, only later.

What’s missing?

While Bausell generally gives us an upbeat and forward-looking discussion, I confess that I miss the sense of excitement that I’ve sometimes felt, for example at a couple of APS Conventions, roughly around 2014-2015. At those, some talks and symposia sparked animated discussions and a real sense that right here we are trying to figure out how science could be done better.

I also have in mind the youth of many advocates of better practices. I’ve had young researchers approach me after talks, sometimes in tears. They know what’s needed, but how, they ask, should they cope with a professor who says “reanalyse—your job is to find statistical significance in your data”? Despite such pressures, the Society for the Improvement of Psychological Science (SIPS) has always been energised by young folks. One-third of its board positions are reserved for graduate students and postdocs. I continue to be inspired and excited by young researchers who are insisting that science is done better, and who are active in helping figure out just how.

My conclusion is that Bausell gives us a valuable, wide-ranging and well-informed history, progress report, and agenda for the Open Science project.

Enjoy,

Geoff

Bausell, R. B. (2021). The Problem With Science: The Reproducibility Crisis and What to Do About It. OUP. ISBN-13: 9780197536537 https://doi.org/10.1093/oso/9780197536537.001.0001

Cardiac Surgery: Yet One More Research Field Highly Critical of p Values

Replacement heart valves, bypasses, transplants: Cardiac surgery research has given us these life-saving goodies, and more. Now this vital research field has joined many others in appreciating the damage that reliance on p values can bring.

Our critical review (here) in the Journal of Cardiac Surgery has just been released online:

David McGiffin, lead author, is a distinguished researcher and cardiac surgeon, specialising in complex transplants. Paul Myers, co-author, is a distinguished researcher and anaesthetist.

David, originally from Queensland, explains that he became increasingly uneasy about p values during his decades of research and clinical practice. “Sometimes, it seemed that researchers can dial up just about any result they wanted.” In other words, p-hacking, cherry-picking, and other Questionable Research Practices seemed to be all around.

In 2013 he took up his present position at Monash. He started reading about p values, came across some of my work, and got in touch. We quickly discovered that we shared many views. David kindly arranged for the invitation that led to me giving two talks at a conference on cardiothoracic surgery in 2018, as I blogged about here. As usual, I found it great fun.

David started work on a p value article for his research field, then involved colleagues and undertook simulations. Several years and after many discussions and revisions, our review emerged.

Our review

In our review, we first discuss the weird backward logic of p values and NHST and how that leads to users being often misled by the inverse probability fallacy—which underlies many p value misconceptions. Then we have a section on the enormous sampling variability of the p value: Replicate an experiment and you are likely to get a very different p value, so p values are highly unreliable, and not to be trusted.

We illustrate that unreliability with the striking figure (below, courtesy of Simon Gandevia) from our reply to a recent editorial by Simon in the Journal of Physiology. My post about that is here.

Lessons from simulations

Then we report three simulation studies run by John Reynolds that used a large database of cardiac surgery cases to explore ways that NHST calculations are typically conducted. We conclude that:

  1. Assumptions matter. For example, if a measure is distinctly not normally distributed in the population, inference calculations can be highly misleading. In some cases a transformation can help, but should be specified in advance.
  2. Failing to reject a null hypothesis is never sufficient justification for accepting that null, and this is especially the case when statistical power is low.
  3. Inference calculations can be misleading when analysing rates, especially when rates are low and some subgroups are very small.
  4. CI and p value calculations can suggest differing conclusions for a number of reasons, including different approximate calculation methods for the two, even if both methods are the defaults in a statistical package.

These conclusions are not new, but the simulations provide striking illustrations in a context familiar to cardiac surgery researchers.

Finally

After some comments—controversial in my view—about cases where p values may have some value, we describe a range of ways that p values can distort science. We mention HARKing—hypothesising after the results are known—and JARKing—justifying after the results are known. ‘JARK’ was new to me, but is a nice acronym. A three number summary, referring to an estimated effect size and the lower and upper limits of the CI on that estimate, is another nice expression particularly familiar to those in the field.

Our conclusion is that cardiac surgery researchers should adopt Open Science practices and use estimation, or other improved approaches, perhaps Bayesian.

Enjoy,

Geoff

McGiffin, D. C., Cumming G., & Myles, P. S. (2021) The frequent insignificance of a “significant” p value. Journal of Cardiac Surgery, 1-10. https://doi.org/10.1111/jocs.15960 Full text is here.

The New APA Style: Try to Contain Your Excitement—and Watch Out for Dud Copies

This is a post for the nerds, fine people that we are. Actually, everyone needs to think about reporting style, especially for statistical stuff.

Seventh edition, 2020
Sixth edition, 2010

The APA Publication Manual is the bible of APA Style, also perhaps the bane of some students’ lives. In ITNS we used the 6th edition (APA, 2010). Now there’s a 7th edition (APA, 2020). I guess we need to switch for ITNS, second edition, tho’ I’m finding it hard to get enthusiastic.

Figures: captions out, headings in

For me, the biggest change is the demise of detailed figure captions placed below the figure. Now there’s a brief heading above a figure, and Note below to give the description. That matches style for tables, but I’m grieving already.

The OLD style for figures, as in ITNS

Figure 1.1. Support for Proposition A, in percent, as reported by the poll. The dot marks the point estimate, and the two lines display the margin of error (2%) either side of the dot. The full interval, from 51% to 55%, is the 95% confidence interval.

The NEW style for figures, as in ITNS2?

Figure 1.1

Support for Proposition A, as Reported by the Poll

Note. Support is expressed in percent. The dot marks the point estimate, and the two lines display the margin of error (2%) either side of the dot. The full interval, from 51% to 55%, is the 95% confidence interval.

What else is new in the 7th edition

A guide to the changes from 6th to 7th editions is here. A video that illustrates the main changes is here. It shows how the new style is a bit simpler—hooray—despite the new edition being more detailed and having more than 50% more pages than the old.

There are small changes to reference formats and how references are cited in text. Reference formats, and examples, are provided for more than 100 types of items, including social media items, blog posts and comments on blog posts, TED talks, songs, and just about every other type of item you can think of.

For the first time, format for student papers is discussed, with great scope given for instructors to specify how student work should be presented. Good.

JARS, the Journal Article Reporting Standards, are highly detailed, and now comprise a whole chapter (Chapter 3). As well as quantitative, they cover qualitative and mixed methods research. There’s particular mention of meta-analyses, and replication studies. Yay.

Only the super-nerds need notice that now only one space is required at the end of a sentence, not two. A massive saving of virtual trees?

Open Science?

There’s brief mention of Open Science badges, but the index includes no other item for ‘Open Science’. There’s no index entry for ‘preregistration’, but the ‘registration’ entry points to a couple of mentions of study registration, in particular of clinical trials. There’s advice on how to include access information about open data and open materials provided online. Overall, however, there is no strong advocacy of Open Science practices. Sadly.

The new Publication Manual: Beware counterfeits

Last year I bought a copy of the 2020 Manual on the Australian site of Amazon. I noticed small errors, for example incorrect italics. I enquired of APA, sent them photos of a few pages, and was told I had a counterfeit copy. It came from a third-party seller via Amazon. I returned the copy to Amazon, as requested, and gave the seller a blistering review—they even had the cheek to ask me to withdraw this as spoiling their business. Happily, APA sent me a replacement copy, although I suspect that’s not usual practice.

There’s now a blog post warning of counterfeits, as well as advice here and here on spotting duds.

Are journals using the new style?

Not surprisingly, APA journals are using the new style. The change is in progress. Skim through recent issues and see a mixture of articles using the old style for figures, and those using the new. For APS journals, the online guidelines for authors still refer to the old 2010 style. I enquired of Elaine Walker, Chair of the APS Publications Committee, who kindly explained that APS intends to swap to the 7th edition, and is currently planning the move. So, I guess Bob and I need to switch. Sigh. I’m already nostalgic for fully explanatory captions under figures.

Why we need reporting standards

Years ago some students and I scanned articles in economics, business, and chemistry journals to see how these disciplines reported statistical significance testing. This was pilot testing, never written up. We found that reporting was often sketchy and obscure. We saw text such as: “Figure 1 illustrates that only groups Groups A and C showed an effect.” It took investigation and guesswork to figure out that the authors had probably used a criterion of 2xSE to identify effects as existing or not. Dichotomous decision making, without even coming clean, stating the criterion being used, and admitting that any difference just less than 2xSE was regarded as not existing.

In Psychology, long-standing APA style requirements have meant that NHST is almost always more full reported. We should be told at least the summary statistics, what’s being tested, the df, the p value, and then a conclusion.

Now that researchers are using computer scanning of very large numbers of journal articles to study, for example, the use of NHST and the distribution of reported p values, it’s more vital than ever that authors report fully and comply with style standards in reporting statistical analyses.

At least to some extent, we all need to embrace our inner nerd. The 2020 version of course.

Geoff

The Journal of Physiology Adopts Better Practices and Excoriates p Values

In brief

The Journal of Physiology is embracing a number of Open Science practices and has just published an editorial highly critical of the p value. Yay!

In less brief

Simon Gandevia, of the NeuRA research institute in Sydney, has been awarded Honorary Membership of The Century Citation Club of JPhysiol, having published more than 100 articles (phew!) in that journal. He was invited to write an editorial, which was published in January. Simon reflected on changes in the Journal and the discipline—massive advances in techniques, larger teams, longer articles, stronger links to clinical applications—and closed by explaining how unreliable and misleading p can be, especially in relation to replications.

Simon’s editorial, statistical aspects

Simon described how JPhysiol had, since around 2010, encouraged authors to adopt better statistical techniques and reporting practices, and had published how-to guides to help. However, little changed, and so in 2018 the Journal mandated more complete reporting of research—to facilitate replication—and especially of data and statistical analyses. In other words, adoption of key Open Science practices.

Simon then used his lovely figure, above, to demonstrate that an Initial p value (p found in an initial study, as shown on the horizontal axis) gives extremely little information about what p an exact replication will find. The solid line (refer to left axis) tells us that initial p = .05 gives only a 50:50 chance that a replication finds p < .05. Initial p = .01 gives only about a 67% chance of replication p < .05.

The vertical grey bars (refer to right axis) are the 80% prediction intervals for replication p. These come from my p intervals article (Cumming, 2008, here or here). If initial p = .05, there’s an 80% chance that p in an exact replication lies in (.0002, .65), and a 10% chance it lies to the left of that interval, and 10% to the right. For initial p = .01, the prediction interval is (.00001, .41). The intervals are so long! A replication can, alas, give just about any p value, so no p is to be trusted.

Simon also discusses the advantages of replacing NHST “with presentation of effect sizes and confidence intervals (ESCI). This… would avoid the phoney dichotomy of significant vs. insignificant…”. Indeed! He mentioned a recent article of his that’s one of a number in JPhysiol in which authors had taken such an estimation approach. That’s progress!

Simon’s editorial is definitely worth a read.

Simon’s editorial: A response and our reply

In reply, Brent Raiteri made a number of points and, in particular, argued that there are typically so many unknowns about a replication that it’s not possible to calculate any value for replication p. Simon kindly invited three of us to join him in a reply to Raiteri, which has just come online. We clarified a couple of points and Simon prepared an amended version of his figure that makes clear that two-tailed p values are used throughout. (It’s this amended figure that I include above.) We also noted that the values in the figure don’t rely on the assumption—unrealistic, but often made—that the initial study estimated exactly the size of the effect in the population.

We noted that the figure assumes that sampling variability is the only cause for differences between initial and replication p, so the probabilities and prediction intervals depicted represent a best case—replication p may, in practice, be even more unreliable. We agreed with Raiteri that usually there are further differences between initial and replication studies, perhaps sometimes sufficient to justify Raiteri’s claim that it’s not possible to calculate replication p values.

The Journal of Physiology

You probably know that JPhysiol is one of the longest-running and most highly regarded journals in the biological sciences. Founded in 1878, it has published classic research from numerous Nobel Laureates, as well as other leading scientists. Many neuroscience courses still ask students to read classic JPhysiol articles, perhaps about the sodium pump, and other fundamental discoveries. It’s especially pleasing to see such a journal adopt and promote Open Science practices.

Research Quality at NeuRA

Within his institute, Simon has long championed improved research practices. The Research Quality page of NeuRA introduces the Reproducibility & Quality Sub-Committee that Simon convenes. It’s active in promoting Open Science practices by NeuRA’s researchers.

At that page, scroll down to see the video of a March 2021 talk by Simon:

Research Quality and Reproducibility: Why You Should Be Worried

  • At about 8.45, note a nice story about Sir John Eccles being acutely aware of having published incorrect findings, then finding a way to do better—which led to his Nobel Prize.
  • At about 22.30, see the Quality Output Checklist and Content Assessment (QuOCCA), an instrument for assessing the transparency, data analysis, and reporting practices of a draft journal manuscript.

A little lower on that page is a video of a talk that I, at Simon’s invitation, gave at NeuRA in December 2019:

Improving the Trustworthiness of Neuroscience Research

I included two demonstrations of the unreliability of replication p values:

  • At about 11.00 see the dance of the p values (or search YouTube for ‘dance of the p values).
  • At about 13.00 I move on to explain then demonstrate significance roulette (or search YouTube for ‘significance roulette’ to find two videos).

The slides of that talk are here and a blog post about my visit to NeuRA is here.

Finally

A warm salute to Simon and the editors at JPhysiol for bringing that august journal into the world of Open Science and better statistics.

Geoff

Reminder: Significance Roulette Still Tells Us a p Value Can’t be Trusted

An appreciative comment on YouTube reminds me I haven’t mentioned Significance Roulette for a while, yet its message that a p value can’t be trusted remains as relevant as ever.

Dance of the p Values

The dance of the p values was my first go at making vivid the amazingly large sampling variability of the p value. The original video is here, or search YouTube for ‘dance of the p values’. A simulation of running the same experiment over and over finds that p values typically leap around wildly: Usually, the next p value can take just about any value.

But that’s when we know the population mean. What about a more realistic situation when all we know from the initial experiment is the p value? What is a close replication, just the same but with a new sample, likely to give? In other words, what’s replication p likely to be?

In most cases replication p can take just about any value 🙁

Significance Roulette

For explanation and all the formulas, see this paper (cited 375 times). For the demo, search YouTube for ‘Significance Roulette’ to find two videos. Or they are here and here.

The figure above is a wheel that’s equivalent to the distribution of replication p following an initial experiment that gives p = .05. (All p values are two-tailed.) Amazingly, the wheel applies whatever the N (unless very small), the power, or the population effect size. All we need is that p = .05, then spinning the wheel is equivalent to running a replication, if all we are interested in is the p value given by that replication. It’s also way cheaper and easier.

Of course, if the initial p value is different we’ll need a different wheel. Below is the wheel for initial p = .01. More red (p < .001) and less deep blue (p > .10), but still an alarmingly wide spread of possibilities.

You don’t believe me? It took me ages to accept what the formulas and simulations were telling me. But they are correct. We really, really can’t trust a p value–which seems to promise certainty, a clear outcome. Far, far better to use the confidence interval, whose length makes the degree of uncertainty salient, even tho’ that’s often a depressing message.

Geoff