The Journal of Physiology Adopts Better Practices and Excoriates p Values

In brief

The Journal of Physiology is embracing a number of Open Science practices and has just published an editorial highly critical of the p value. Yay!

In less brief

Simon Gandevia, of the NeuRA research institute in Sydney, has been awarded Honorary Membership of The Century Citation Club of JPhysiol, having published more than 100 articles (phew!) in that journal. He was invited to write an editorial, which was published in January. Simon reflected on changes in the Journal and the discipline—massive advances in techniques, larger teams, longer articles, stronger links to clinical applications—and closed by explaining how unreliable and misleading p can be, especially in relation to replications.

Simon’s editorial, statistical aspects

Simon described how JPhysiol had, since around 2010, encouraged authors to adopt better statistical techniques and reporting practices, and had published how-to guides to help. However, little changed, and so in 2018 the Journal mandated more complete reporting of research—to facilitate replication—and especially of data and statistical analyses. In other words, adoption of key Open Science practices.

Simon then used his lovely figure, above, to demonstrate that an Initial p value (p found in an initial study, as shown on the horizontal axis) gives extremely little information about what p an exact replication will find. The solid line (refer to left axis) tells us that initial p = .05 gives only a 50:50 chance that a replication finds p < .05. Initial p = .01 gives only about a 67% chance of replication p < .05.

The vertical grey bars (refer to right axis) are the 80% prediction intervals for replication p. These come from my p intervals article (Cumming, 2008, here or here). If initial p = .05, there’s an 80% chance that p in an exact replication lies in (.0002, .65), and a 10% chance it lies to the left of that interval, and 10% to the right. For initial p = .01, the prediction interval is (.00001, .41). The intervals are so long! A replication can, alas, give just about any p value, so no p is to be trusted.

Simon also discusses the advantages of replacing NHST “with presentation of effect sizes and confidence intervals (ESCI). This… would avoid the phoney dichotomy of significant vs. insignificant…”. Indeed! He mentioned a recent article of his that’s one of a number in JPhysiol in which authors had taken such an estimation approach. That’s progress!

Simon’s editorial is definitely worth a read.

Simon’s editorial: A response and our reply

In reply, Brent Raiteri made a number of points and, in particular, argued that there are typically so many unknowns about a replication that it’s not possible to calculate any value for replication p. Simon kindly invited three of us to join him in a reply to Raiteri, which has just come online. We clarified a couple of points and Simon prepared an amended version of his figure that makes clear that two-tailed p values are used throughout. (It’s this amended figure that I include above.) We also noted that the values in the figure don’t rely on the assumption—unrealistic, but often made—that the initial study estimated exactly the size of the effect in the population.

We noted that the figure assumes that sampling variability is the only cause for differences between initial and replication p, so the probabilities and prediction intervals depicted represent a best case—replication p may, in practice, be even more unreliable. We agreed with Raiteri that usually there are further differences between initial and replication studies, perhaps sometimes sufficient to justify Raiteri’s claim that it’s not possible to calculate replication p values.

The Journal of Physiology

You probably know that JPhysiol is one of the longest-running and most highly regarded journals in the biological sciences. Founded in 1878, it has published classic research from numerous Nobel Laureates, as well as other leading scientists. Many neuroscience courses still ask students to read classic JPhysiol articles, perhaps about the sodium pump, and other fundamental discoveries. It’s especially pleasing to see such a journal adopt and promote Open Science practices.

Research Quality at NeuRA

Within his institute, Simon has long championed improved research practices. The Research Quality page of NeuRA introduces the Reproducibility & Quality Sub-Committee that Simon convenes. It’s active in promoting Open Science practices by NeuRA’s researchers.

At that page, scroll down to see the video of a March 2021 talk by Simon:

Research Quality and Reproducibility: Why You Should Be Worried

  • At about 8.45, note a nice story about Sir John Eccles being acutely aware of having published incorrect findings, then finding a way to do better—which led to his Nobel Prize.
  • At about 22.30, see the Quality Output Checklist and Content Assessment (QuOCCA), an instrument for assessing the transparency, data analysis, and reporting practices of a draft journal manuscript.

A little lower on that page is a video of a talk that I, at Simon’s invitation, gave at NeuRA in December 2019:

Improving the Trustworthiness of Neuroscience Research

I included two demonstrations of the unreliability of replication p values:

  • At about 11.00 see the dance of the p values (or search YouTube for ‘dance of the p values).
  • At about 13.00 I move on to explain then demonstrate significance roulette (or search YouTube for ‘significance roulette’ to find two videos).

The slides of that talk are here and a blog post about my visit to NeuRA is here.

Finally

A warm salute to Simon and the editors at JPhysiol for bringing that august journal into the world of Open Science and better statistics.

Geoff

Reminder: Significance Roulette Still Tells Us a p Value Can’t be Trusted

An appreciative comment on YouTube reminds me I haven’t mentioned Significance Roulette for a while, yet its message that a p value can’t be trusted remains as relevant as ever.

Dance of the p Values

The dance of the p values was my first go at making vivid the amazingly large sampling variability of the p value. The original video is here, or search YouTube for ‘dance of the p values’. A simulation of running the same experiment over and over finds that p values typically leap around wildly: Usually, the next p value can take just about any value.

But that’s when we know the population mean. What about a more realistic situation when all we know from the initial experiment is the p value? What is a close replication, just the same but with a new sample, likely to give? In other words, what’s replication p likely to be?

In most cases replication p can take just about any value 🙁

Significance Roulette

For explanation and all the formulas, see this paper (cited 375 times). For the demo, search YouTube for ‘Significance Roulette’ to find two videos. Or they are here and here.

The figure above is a wheel that’s equivalent to the distribution of replication p following an initial experiment that gives p = .05. (All p values are two-tailed.) Amazingly, the wheel applies whatever the N (unless very small), the power, or the population effect size. All we need is that p = .05, then spinning the wheel is equivalent to running a replication, if all we are interested in is the p value given by that replication. It’s also way cheaper and easier.

Of course, if the initial p value is different we’ll need a different wheel. Below is the wheel for initial p = .01. More red (p < .001) and less deep blue (p > .10), but still an alarmingly wide spread of possibilities.

You don’t believe me? It took me ages to accept what the formulas and simulations were telling me. But they are correct. We really, really can’t trust a p value–which seems to promise certainty, a clear outcome. Far, far better to use the confidence interval, whose length makes the degree of uncertainty salient, even tho’ that’s often a depressing message.

Geoff

Fiona’s Take on Richie’s New Book: ‘Science Fictions–How Fraud, Bias, Negligence, and Hype…’

I haven’t read this book–it won’t reach here for a while–but Fiona has written a critique that’s certainly worth reading. Out this week in Nature, you can find her assessment here.

From Fiona’s review:

“Together with his overview of the replication crisis, this introduction would be useful for undergraduates or general readers.

“Fraud, bias, negligence and hype are the themes of Science Fictions. Some of the cases Ritchie presents… are intriguing and disturbing combinations of all four. 

“This comprehensive collection of mishaps, misdeeds and tales of caution is the great strength of Ritchie’s offering.” 

Then come the really interesting bits, including discussion of Richie’s interpretation of the problem and his view of what science should be and what he sees it as having become. Fiona can expertly set that all in perspective, as a scholar of the history and philosophy of science, and now metascience. Definitely worth the read. She’d recommend the book also, despite some flaws.

Geoff

Richie, S. (2020). Science fictions: How fraud, bias, incompetence, and hype undermine the search for truth. Metropolitan Books.

Will that result replicate? Contribute some judgments to Fiona’s repliCATS project

Just had an update about Fiona’s big project to study how well researchers can judge the chance that some result can be replicated. She and the team are doing well, but need a last push during July to meet their target. Consider signing up to make a few judgments (and maybe win some cash). Sorry–I’m not really a used-car salesperson.

I posted about the project, repliCATS, here. In short: “…the largest ever empirical study on how scientists reason about other scientists’ work, and what factors makes them trust it.” It’s right at the core of understanding and advancing Open Science.

At workshops around the world, more recently all online, the team has collected a big database of judgments about claims made in reports of research. Would the claim replicate? The team seeks judgments by anyone from undergraduates to seasoned researchers, in any of a wide range of social and behavioural science disciplines.

Please do consider signing up. Details are here.

Latest news: The project has just been expanded to include assessment of claims being made in the social and behavioural sciences about COVID-19. Work to start in August. You can sign up for that also.

eNeuro Keeps Up the Good Work on Estimation

Update: The great figures below were produced using the wonderful Estimation Statistics package. See also Ho, et al. (2019).

Back in August Bob posted (here) about eNeuro‘s great initiative to encourage authors to use estimation. The latest eNeuro email update reports how the journal is keeping up the good work. See three paras below:

eNeuro encourages authors to add estimation statistics to their analyses when appropriate. Below, we will feature papers that did so, along with an author’s response to the query: What value was added to your analysis or perspective through the addition of estimation statistics? For more information read Estimation for Better Inference in Neuroscience.

Circuit-specific dendritic development in the piriform cortex
Laura Moreno-Velasquez, Hung Lo, Stephen Lenzi, Malte Kaehne, Jörg Breustedt, Dietmar Schmitz, Sten Rüdiger, and Friedrich W. Johenning

“It was satisfying to see how switching to estimation stats based data display increased the transparency of our data display in a reader-friendly format. By openly displaying the mean effect sizes and confidence intervals, we felt relieved from the pressure to solely rely on p-value based true/false statements about the data. Estimation statistics also makes it obvious to us and our readers where replications and further experiments are most needed in the future.” — Friedrich W. Johenning

—that’s all great to see. Below are a couple of small parts of the figures that the author refers to:

The lower means and CIs, with plausibility curves, illustrate the estimated differences between the means of the blue dots (L2B cells, whatever they are) and red dots (L2A cells). The author mentioned ‘transparency’; indeed, the pictures do make it easy to appreciate the differences, and the precision with which each was estimated.

I especially love the author’s final sentence:

Estimation statistics also makes it obvious to us and our readers where replications and further experiments are most needed in the future.

Science as a cumulative, progressive enterprise–powered by estimation. Yay!

Geoff

Ho, J., Tumkaya, T., Aryal, S. et al. Moving beyond P values: data analysis with estimation graphics. Nat Methods 16, 565–566 (2019). https://doi.org/10.1038/s41592-019-0470-3

JAMA Opposes p<.05 Decision Making

A recent viewpoint article (free download) in JAMA argues that decisions should not be based on mere p value criteria, but need “consideration of the outcome in terms of effect size and accompanying CIs, [and] placing the findings from the trial in the context of the totality of the existing relevant evidence”. Hooray for JAMA! The overview:

JAMA. Published online May 8, 2020. doi:10.1001/jama.2020.3508 

A common argument given in defence of NHST is that (1) in the real world we need decisions, (2) NHST and p<.05 is a way to make such clear decisions that is (3) based on the evidence and (4) objective. Yes, 1 and 2 are true, but 3 is only partly true, and 4 may be true if all details of the data analysis and decision procedure are preregistered. However, a p value reflects N as much as anything, and, most importantly, other vital factors need to be considered in making a real-world decision, beyond the current data. Decisions need to be informed by data, of course, preferably via confidence intervals and meta-analysis. But they also should reflect consideration of alternatives, costs and benefits, the values of relevant parties, and so on. ‘Conclusions’ above acknowledges all this. Hooray for JAMA!

That’s all fine, but the title and main consideration of the viewpoint are surprising: What about non-statistically significant results? Given the recommended approach to decision making, the viewpoint argues that there may even be cases in which such a non-sig result might justify a change in clinical practice. Phew–it seems a very forced argument to me: Take some very weak evidence and dream up a case in which that might just tip the balance and lead us to change practice? Perhaps, but …

The examples discussed, however, emphasise the role of prior evidence (meta-analysis is not mentioned explicitly) and of costs and benefits. In addition, the extent of uncertainty, even if a clinical decision has to be made, needs emphasis. So that’s all good.

The article includes this link to an interview by Howard Bauchner, JAMA’s Editor in Chief, with Paul Young, author of the viewpoint. OK, it’s quicker to skim the article than to listen to the interview, but the interview makes clear that Bauchner takes seriously the need to move on from p<.05. That, for me, is the main point. Hooray for JAMA!

Geoff

P.S. Many thanks to Anoop Balachandran for the heads up.

 

Teaching in the New Era of Psychological Science

A great collection of articles in the latest issue of PLAT. Contents page here. At that page, click to see the abstract of any article.

A big shout out to the wonderful editorial team that assembled this special issue: Susan A. Nolan (Seton Hall University, USA), Tamarah Smith (Gwynedd Mercy University, USA), Kelly M. Goedert (Seton Hall University, USA), Robert Calin-Jageman (Dominican University, USA)

The intro (as above) gives a good overview and summary of all the articles. It’s on open access here. As the editors say, the issue includes two reviews, three articles, and three reports, which together cover a wide range of Open Science issues, all from a teaching perspective.

Here are my haphazard thoughts:

  • Anglin and Edlund report a survey of psychology instructors. Overall, most don’t teach much about OS practices, but believe that more should be done. They identify the current incentive structure for researchers as a big problem and high priority for change. Indeed!
  • Three articles include discussion of Bayesian approaches to data analysis and interpretation. It’s certainly a big issue how statistical inference practice should and will develop in psychology. Can beginners be introduced successfully to Bayesian methods? (Play with JASP  and read the van Doorn et al. article for an inkling of how things are developing.) Should students be exposed to both conventional and Bayesian approaches to statistical thinking, or will this merely create confusion? If so, at which stage in a student’s statistical education? If first one then the other, which should come first?
  • I’m struck by the extent of student involvement and activity in many of the courses and approaches described. Flipped classroom, student-led discussions and classes, expectations of student initiative, and more. Very heartening! Such approaches tie in well with projects that follow the full sequence of a research project, from question conception through all the steps to full reporting and contemplation of further replication. I wish I, way back as a student, had experienced such courses!
  • Several articles describe courses and projects in which students collaborate, perhaps across several universities, to carry out a replication of a published study. Involving groups of students means that the replication can be usefully large. Following OS practices means that the results should be publishable, and a useful contribution to the literature. Such projects can also provide an excellent educational experience for upper year undergraduates and perhaps masters students.
  • Back in 2015 at the WPA Annual Convention in Las Vegas I was part of a symposium on collaborative replication projects involving students. Many of those attending the symposium were faculty in liberal arts schools, many of whom had little scope to conduct research, but who were expected to train their students in research methods. I felt a palpable enthusiasm in the room for this very new (at that time) idea that their students could participate in worthwhile large-scale collaborative replication projects that could provide excellent training, while being practically achievable within the limited resources available in many cases. Several of the PLAT articles describe how this approach has now developed considerably.
  • Taking a broad perspective, the editors “wonder if the approach these authors outlined might also be a way to combat an anti-science climate. If a greater number of students are actually engaged in science (e.g., in replications) rather than just class projects, they may view themselves as part of a larger scientific community, and, in turn, be less likely to have an overall distrust of science.” Bring this on!

Enjoy,

Geoff

 

 

 

 

 

 

 

 

 

 

Farewell and Thanks Steve Lindsay

Psychological Science, the journal, has for years pushed hard for publication of better, more trustworthy research. First there was the leadership of Eric Eich, then Steve Lindsay energetically took the baton. Steve is about to finish, no doubt to his great relief. His ‘swan song’ editorial has just come online, with open access:

Steve starts with a generous reference to a talk of mine at Victoria University in Wellington N.Z. No doubt I talked mainly about the huge variability of p values, and ran the ‘dance of the p values‘ demo. (Search for ‘dance of the p values’ at YouTube to find maybe 3 versions.)

Then he gives a brief and modest account of the great strides the journal took under Eric’s and then his own leadership. Indeed, Psychological Science has been vitally important in advancing our discipline! I’m sure, also, that it has had beneficial influence well beyond psychology.

All eyes are now on the incoming editor, Patricia Bauer. To what extent will she keep up the Open Science pressure, the further development of journal policies and practices to keep our field moving towards ever more reproducible and open–and therefore trustworthy and valuable–research?

I join Steve in wholeheartedly wishing her well.

Geoff

NeuRA Ahead of the Open Science Curve

I had great fun yesterday visiting NeuRA (Neuroscience Research Australia), a large research institute in Sydney. I was hosted by Simon Gandevia, Deputy Director, who has been a long-time proponent of Open Science and The New Statistics.

Neura’s Research Quality page describes the quality goals they have adopted, at the initiative of Simon and the Reproducibility & Quality Sub-Committee, which he leads. Not only goals, but strategies to bring their research colleagues on board–and to improve the reproducibility of NeuRA’s output. My day started with a discussion with this group. They described a whole range of projects they are working on to strengthen research at NeuRA–and to assess how quality is (they hope!) increasing.

For example, Martin Heroux described the Quality Output Checklist and Content Assessment (QuOCCA) tool that they have developed, and are now applying to recent past research publications from NeuRA. In coming years they plan to assess future publications similarly–so they can document the rapid improvement!

I should mention that Martin and Joanna Diong run a wonderful blog, titled Scientifically Sound–Reproducible Research in the Digital Age.

It was clear that the R&Q folks, at least, were very familiar with Open Science issues. Would my talk be of sufficient interest for them? Its title was Improving the trustworthiness of neuroscience research (it should have been given by Bob!), and the slides are here. The quality of the questions and discussion reassured me that at least many of the folks in the audience were (a) on board, but also (b) very interested in the challenges of Open Science.

After lunch my ‘workshop’ was actually a lively roundtable discussion, in which I sometimes managed to explain a bit more about Significance Roulette (videos are here and here), demo a bit of Bob’s new esci in R, or join in brainstorming strategies for researchers determined to do better. My slides are here.

Yes, great fun for me, and NeuRA impresses as working hard to achieve reproducible research. Exactly what the world needs.

Geoff

Replications: How Should We Analyze the Results?

Does This Effect Replicate?

It seems almost irresistible to think in terms of such a dichotomous question! We seem to crave an ‘it-did’ or ‘it-didn’t’ answer! However, rarely if ever is a bald yes-no decision the most informative way to think about replication.

One of the first large studies in psychology to grapple with the analysis of replications was the classic RP:P (Replication Project: Psychology) reported by Nosek and many colleagues in Open Science Collaboration (2015). The project identified 100 published studies in social and cognitive psychology then conducted a high-powered preregistered replication of each, trying hard to make each replication as close as practical to the original.

Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349 (6251) aac4716-1 to aac4716-8.

The authors discussed the analysis challenges they faced. They reported 5 assessments of the 100 replications:

  • in terms of statistical significance (p<.05) of the replication
  • in terms of whether the 95% CI in the replication included the point estimate in the original study
  • comparison of the original and replication effect sizes
  • meta-analytic combination of the original and replication effect sizes
  • subjective assessment by the researchers of each replication study: “Did the effect replicate?”

Several further approaches to analyzing the 100 pairs of studies have since been published.

Even so, the one-liner that appeared in the media and swept the consciousness of psychologists was that RP:P found that fewer than half of the effects replicated: Dichotomous classification of replications rules! For me, the telling overall finding was that the replication effect sizes were, on average, just half the original effect sizes, with large spreads over the 100 effects. This strongly suggests reporting bias, p-hacking or some other selection bias influenced some unknown proportion of the 100 original articles.

OK, how should replications be analyzed? Happily, there has been progress.

Meta-Analytic Approaches to Assessing Replications

Larry Hedges, one of the giants of meta-analysis since the 1980s, and Jacob Schauer recently published a discussion of meta-analytic approaches to analyzing replications:

Hedges, L. V., & Schauer, J. M. (2019). Statistical analyses for studying replication: Meta-analytic perspectives. Psychological Methods, 24, 557-570. http://dx.doi.org/10.1037/met0000189

Abstract
Formal empirical assessments of replication have recently become more prominent in several areas of science, including psychology. These assessments have used different statistical approaches to determine if a finding has been replicated. The purpose of this article is to provide several alternative conceptual frameworks that lead to different statistical analyses to test hypotheses about replication. All of these analyses are based on statistical methods used in meta-analysis. The differences among the methods described involve whether the burden of proof is placed on replication or nonreplication, whether replication is exact or allows for a small amount of “negligible heterogeneity,” and whether the studies observed are assumed to be fixed (constituting the entire body of relevant evidence) or are a sample from a universe of possibly relevant studies. The statistical power of each of these tests is computed and shown to be low in many cases, raising issues of the interpretability of tests for replication.

The discussion is a bit complex, but here are some issues that struck me:

  • We usually wouldn’t expect the underlying effect to be identical in any two studies. What small difference would we regard as not of practical importance? There are conventions, differing across disciplines, but it’s a matter for informed judgment. In other words, how different could underlying effects be, while still justifying a conclusion of successful replication?
  • Should we choose fixed-effect or random-effects models? Random-effects is usually more realistic, and what ITNS and many other books recommend for routine use. However, H&S use fixed-effect models throughout, to limit the complexity. They report that, for modest amounts of heterogeneity their results do not differ greatly from what random-effects would give.
  • Meta-analysis and effect size estimation are the focus throughout, but even so the main aim is to carry out a hypothesis test. The researcher needs to choose whether to place the burden of proof on nonreplication or replication. In other words, is the null hypothesis that the effect replicates, or that it doesn’t?
  • One main conclusion from the H&S discussion and the application of their methods to psychology examples is that replication projects typically need many studies (often 40+) to achieve adequate power for those hypothesis tests, and that even large psychology examples are under-powered.

A further sign that the issues are complex is the comment published immediately following H&S with suggestions for an alternative way to think about heterogeneity and replication:

Mathur, M. B., & VanderWeele, T. J. (2019). Challenges and suggestions for defining replication “success” when effects may be heterogeneous: Comment on Hedges and Schauer (2019). Psychological Methods, 24, 571-575. http://dx.doi.org/10.1037/met0000223

H&S gave a brief reply:

Hedges, L. V., & Schauer, J. M. (2019). Consistency of effects is important in replication: Rejoinder to Mathur and VanderWeele (2019). Psychological Methods, 24, 576-577. http://dx.doi.org/10.1037/met0000237

Our Simpler Approach

I welcome the above three articles and would study them in detail before setting out to design or analyze a large replication project.

In the meantime I’m happy to stick with the simpler estimation and meta-analysis approach of ITNS.

Meta-Analysis to Increase Precision

Given two or more studies that you judge to be sufficiently comparable, in particular by addressing more-or-less the same research question, then use random-effects meta-analysis to combine the estimates given by the studies. Almost certainly, you’ll find a more precise estimate of the effect most relevant for answering your research question.

Estimating a Difference

If you have an original study and a set of replication studies, you could consider (1) meta-analysis to combine evidence from the replication studies, then (2) finding the difference (with CI of course) between the point estimate found by the original study and that given by the meta-analysis. Interpret that difference and CI as you assess the extent to which the replication studies may or may not agree with the original study.

If the original study was possibly subject to publication or other bias, and the replication studies were all preregistered and conducted in accord with Open Science principles, then a substantial difference would provide evidence for such biases in the original study–although other causes couldn’t be ruled out.

Moderation Analysis

Following meta-analysis, consider moderation analysis, especially if DR, the diamond ratio, is more than around 1.3 and if you can identify a likely moderating variable. Below is our example from ITNS in which we assess 6 original studies (in red) and Bob’s two preregistered replications (blue). Lab (red vs. blue) is a possible dichotomous moderator. The difference and its CI suggest publication or other bias may have influenced the red results, although other moderators (perhaps that red studies were conducted in Germany, blue in the U.S.) may help account for the red-blue difference.

Figure 9.8. A subsets analysis of 6 Damisch and 2 Calin studies. The difference between the two subset means is shown on the difference axis at the bottom, with its CI. From d subsets.

My overall conclusions are (1) it’s great to see that meta-analytic techniques continue to develop, and (2) our new-statistics approach in ITNS continues to look attractive.

Geoff