Meta-Science: It’s all Happening in Melbourne

Are you interested in meta-science? In Open Science? If so, check out the inaugural conference of AIMOS, the Association for Interdisciplinary Research &Open Science. It’s a two-day meeting, on 7 & 8 November, at the University of Melbourne.

There’s an impressive list of confirmed speakers: Click here and scroll down. All the folks I know on that list will definitely be worth hearing.

Here’s how Fiona Fidler and her organising team describe the conference:

What to expect

AIMOS2019 will be a partially unstructured conference. Each of the two days will have a theme, and will start with a series of keynotes or shorter, what we are calling “mini-notes”, followed by a more unstructured part of the day. Check out the draft program here!

We aim for AIMOS2019 to appeal to students and researchers from a range of disciplines with a shared interest in understanding and addressing challenges to replicability, reproducibility and open science. AIMOS2019 will cover a broad range of open science and scientific reform topics, including: pre-registration and Registered Reports; peer review and scientific publishing; using R for analysis; open source experimental programming; meta-research; replicability; improving statistical and scientific inference; diversity in scientific community and practice; and methodological and scientific culture change.

Got questions? You can email the team at aimos-conference@unimelb.edu.au

Make a submission

The call for workshops, unconferences, hackathons and posters is now open. Submit your proposal here.
For further details see here.

Originally, submissions were due by 1 Sep–this Sunday, and beware that Sunday happens earlier down here, maybe approaching 24 hours before you have your Sunday! But now the submission guidelines state that the organisers will consider submissions made up to conference time, and even during the conference, but please get submissions in early to secure a place in the program. I know that already they have received an impressive list of strong proposals, so don’t delay!

Travel grants

There are 55 travel grants of USD400 available for people living far from Melbourne and who are willing to take part in a one-day Replicats workshop on 6 November. For details, go here and scroll down. These workshops have been found absorbing by participants, and contribute mightily to the repliCATS project.

Registration

Registration is open now and is not expensive. The link is here.

Launch of AIMOS

There will be a networking event on Thurs 7 November that will include the formal launch of the Association for Interdisciplinary Meta-Research & Open Science.

Estimation workshop proposal

Bob and I have proposed a workshop on estimation. If accepted, it will include a first glimpse of the R goodies that Bob is developing with the aim of moving ITNS into the R age. Exciting!

November is a great time to visit Melbourne. I hope to see you then!

Geoff

Judging Replicability: Fiona’s repliCATS Project

Judging Replicability

Whenever we read a research article we almost certainly form a judgment of its believability. To what extent is it plausible? To what extent could it be replicated? What are the chances that the findings are true?

What features of the article drive our judgment? Likely influences include:

  • A priori plausibility. To what extent is it unsurprising, reasonable?
  • How large are the effects? How meaningful?
  • The reputation of the authors, the standing of their lab and institution.
  • The reputation of the journal.
  • Interpretations, judgments, claims and conclusions made by the authors.

Open Science ideas will emphasise our attention to features of the research itself, including:

  • Are there multiple lines of converging evidence? Replications?
  • Was the research, including the data analysis strategy, preregistered?
  • Sample sizes? Quality of manipulations (IVs) and measures (DVs)?
  • Results of statistical inference, especially precision (CI lengths)?
  • Are we told that the research was reported in full detail? Are we assured that all relevant studies and results have been reported?
  • Any signs of cherry-picking or p hacking?
  • How prone is the general research area to publish non-replicable results?

Automating the Assessment of Replicability

The SCORE project is a large DARPA attempt to find automated ways to assess the replicability of social and behavioural science research. As I understand it, teams around the world are just beginning on:

  1. Running replications of a large number of published studies, to provide empirical evidence of replicability, and a reference database of studies.
  2. Studying how human experts judge replicability of reported research–how well do they do, and what features (as in the lists above) guide their judgments?
  3. Building AI systems to take the results from (2) and make automated assessments of the replicability of published research.

Brian Nosek, of COS, is leading a project on 1. above. Fiona Fidler is leading one of the projects tackling 2. above: the repliCATS project.

repliCATS

It’s big, maybe up to US$6.5M. It’s ambitious. And Fiona has multiple teams working on various aspects, all with impossibly tight time lines.

A month or so ago I spent a fascinating 3 days down at the University of Melbourne, for research meetings, seminars, and more, as the teams worked on their plans. Brian Nosek was in town, giving great presentations, and consulting as to how his team and Fiona’s could best work together. Here’s the outline of repliCATS, from the project website:

  1. The repliCATS project aims to develop more accurate, better-calibrated techniques to elicit expert assessment of the reliability of social science research.
  2. Our approach is to adapt and deploy the IDEA protocol developed here at the University of Melbourne to elicit group judgements for the likely replicability of 3,000 research claims.
  3. The research we will undertake as part of the repliCATS project will include the largest ever empirical study on how scientists reason about other scientists’ work, and what factors makes them trust it.
  4. We are building a custom online platform to deploy the IDEA protocol. This platform will have a life beyond the repliCATS project: it will be able to be used in the future to enhance expert group judgements on a wide range of topics, in a number of disciplines.

If you are interested in repliCATS, subscribe here for updates. You can also follow Fiona @fidlerfm

I’m agog to follow how it goes, and to see the insights I’m sure they will find into researchers’ judgments about replicability.

Geoff

Joining the fractious debate over how to do science best

At the end of the month (March 2019) the American Statistical Association will publish a special issue on statistical inference “after p values”. The goal of the issue is to focus on the statistical “dos” rather than statistical “don’ts”. Across these articles there are some common themes, but also some pretty sharp disagreements about how best to proceed. Moreover, there is some very strong disagreement about the whole notion of bashing p values and the wisdom of the ASA putting together this special issue (see here, for example).

Fractious argument is the norm in the world of statistical inference, hence the old joke that the plural of “statistician” is a “quarrel”. And why not? Debates about statistical inference get to the heart of epistemology and the philosophy of science–they represent the ongoing normative struggle to articulate how to do science best. Sharp disagreement over the nature of science is the norm–it has always been part of the scientific enterprise and it always will be. It is this intense conversation that has helped define and defend the boundaries of science.

Geoff has long been involved in debates over statistical inference and how to do science best, but this is new to me (Bob). I’m proud of the contribution we submitted to the ASA–I think it’s the best piece I’ve ever written. But I have to say that I go into the debate over inference (and science in general) with some trepidation. First, it is intrinsically gutsy to think you have something to say about how to do science best. Second, I’m the smallest of small-frys in the world of neuroscience–so it’s not like I have notable success at doing science to point to as a support for my claims. Finally, this ongoing debate has a long history and is populated by giants I look up to, most of whom (unlike me) have specialized in studying these topics. In my case, I’ve been learning on the go for the past ten years or so, starting from a foundation that involved plenty of graduate-level stats, but which didn’t even equip me to properly understand the difference between Bayesian and frequentist approaches to statistics.

As I wade into this fraught debate, I thought it might help me to reflect a bit on my own meta-epistemology–to articulate some basic premises that I hold to in terms of thinking about how to fruitfully engage in debate over inference and the philosophy of science. These premises are not only my operating rules, but also my philosophical courage–they explain why I think a noob like me can and should be part of the debate, and why I encourage more of my colleagues in the neurosciences and psychological sciences to tune in and jump in.

There are no knock-out punches in philosophy. This comes from one of my amazing philosophy mentors, Gene Cline. It has taken me a long time to both understand and embrace what he meant. As a young undergrad philosophy major I was eager to demolish–to embarrass Descartes’ naive dualism, to rain hell on Chalmer’s supposedly hard problems of consciousness, and to expose the circular bloviation of Kant’s claims about the categorical imperative. Gene (gradually) helped me understand, though, that if you can’t see any sense in someone’s philosophical position then you’re probably not engaging thoughtfully with their ideas, concerns, or premises (cf Eco’s Island of the Day Before). It’s easy to dismiss straw-person or exaggerated versions of someone’s position, but if you interpret generously and take seriously their best arguments, you’ll find that no deep philosophical debate is easily settled. I initially found this infuriating, but I’ve come embrace it. So I now look with healthy skepticism at those who offer knock-out punches (e.g. (Morey, Hoekstra, Rouder, Lee, & Wagenmakers, 2015)). I hope that in discussing my ideas with others to a) take their claims and concerns seriously, taking on the best possible argument for their position, and b) not to offer my criticisms as a sure and damning refutation… these only seem to exist when we’re not really listening to each other.i

Inference works great until it doesn’t. As Hume pointed out long ago, there is no logical glue holding together the inference engine. Inference assumes that the past will be a good guide to the future, but there is no external basis for this premise, nor could there be (A rare knockout punch in philosophy? Well, even this is still debated). Even if we don’t mind the circularity of induction, we still have to respect the fact that past is not always prelude: inference works great, until it doesn’t (c.f. Mark Twain’s amazing discussion in Life on the Mississippi). So whatever system of inference we want to support we should be clear-eyed that it will be imperfect and subject to error, and that when/how it breaks down will not always be predictable. This is really important in terms of how we evaluate different approaches to statistical inference–none will be perfect under all circumstances, so evaluations must proceed in terms of strengths/weaknesses and boundary conditions. The fact that an approach works poorly in one circumstance is not always a reason to condemn it. We can thoughtfully make use of tools that in under some circumstances are dangerous.

We don’t all want the same things. Science is diverse and we’re not all playing the game in exactly the same way or for the same ends. I see this every year on the floor of the Society for Neuroscience conference, where over 30,000 neuroscientists meet to discuss their latest research. The scope of the enterprise is hard to imagine, and the diversity in terms of what people are trying to do is staggering. That’s ok. We can still have boundaries between science and pseudoscience without having complete homogeneity of statistical, inferential, and scientific approaches. So beware of people telling you what you, as a scientist, want to know. Beware of someone condemning all use of a statistical approach because it doesn’t tell them what they,want to know. That’s my take on a good blog post by Daniel Lakens.

Nullius in verba* Ok – so we have to tread cautiously. But that does not devolve us into sophomoric inferential relativism (everyone’s right in some way; trophies for all!). We can still make distinctions and recognize differences. How? Well, to the extent that there is any “ground truth” in science it is the ability to establish procedures for reliably observing an effect. We could be wrong about what the effect means. But we’re not doing science if we can’t produce procedures that others can use to verify our observations. This is embodied in the founding of the Royal Society, which selected the motto Nullius in verba (verbum), which means “take no one’s word for it” or “see for yourself” (hat tip to a fantastic presentation by Cristobal Young on this). We can evaluate scientific fields for their ability to be generative this way–to establish effects that can be reliably observed and then dissected (not so fast, Psi research). We can also evaluate systems of inference in this way–for their ability (predicted or actual) to help scientists develop procedures to reliably observe effects. By this yardstick some methods of inference will be demonstrably bad (conducting noisy studies and then publishing the statistically significant results as fact while discarding the rest—bad!). But we should expect there to be multiple reasonable approaches to inference, as well as constant space for potential improvement (though usually with other tradeoffs). Oh yeah–this is a very slippery yardstick. It is not easy to discern or predict the fruitfulness of an inferential approach, and there can be strong disagreement about what counts as reliably establishing an effect.

This emphasis on replicability as essential to science cuts a tiny bit against my above point that not all scientists want the same thing. Moreover, in the negative reaction to the replication crisis, I’ve seen some commentaries where there seems to be little concern or regard for the standard of establishing verifiable effects. This, to my mind, stretches scientific pluralism past the breaking point: if you’re not bothered by a lack of replicability of your research, you’re not interested in science.

Authority will only get you so far. The debate over inference has a long history. It’s important not to ignore that . But it is equally important not to use historical knowledge as a cudgel; appeals to authority are not a substitute for good argument. Maybe it is my outside perception, but I feel like quotes from Fisher or Jeffreys or Meehl or sometimes weaponized to end discussion rather than contribute to it.

Ok – so those are my current ideas for how to approach arguments about science and statistical inference: a) embrace real statistical pluralism without letting go of norms and evaluation; b) ground evaluation (as much as possible) in what we think can best foster generative (reproducible) research, c) listen and take the best of what others have to offer, and d) try not to lean too heavily on the Fisher quotes.

At the moment, I’ve landed on estimation as the best approach for the statistical issues I face. I’m confident enough in that choice that I feel good advocating for the use of estimation for other scientists with similar goals. In advocating for estimation, I’m not going to claim a knock-out punch against p values or other approaches, or that the goals estimation can help with are the only legitimate goals to have. Moreover, in advocating for estimation, my goal is not hegemony. Hegemony of misusing p values is where we are currently at, and we don’t need to replace one imperial rule with another. I am helping a journal re-orient its author guidelines towards estimation (with or in place of p values)—but my goal is a diverse landscape of publication options in neuroscience, one where there are outlets for different but fruitful approaches to inference.

Ok – those are my thoughts for now on how to fruitfully debate about statistical inference.  I’m sure I have a lot to learn.  I’m looking forward to the special issue that will soon be out from the ASA and the debate that will surely ensue. 

*Thanks to Boris Barbour for pointing out I misquoted the Royal Society Motto in the original post.

  1. Morey, R. D., Hoekstra, R., Rouder, J. N., Lee, M. D., & Wagenmakers, E.-J. (2015). The fallacy of placing confidence in confidence intervals. Psychonomic Bulletin & Review, 103–123. doi:10.3758/s13423-015-0947-8

Statistical Cognition: An Invitation

Statistical Cognition (SC) is the study of how people understand–or, quite often, misunderstand–statistical concepts or presentations. Is it better to report results using numbers, or graphs? Are confidence intervals (CIs) appreciated better if shown as error bars in a graph or as numerical values?

And so on. These are all SC questions. For statistical practice to be evidence-based, we need answers to SC questions, and these should inform how we teach, discuss, and practise statistics. Of course.

An SC Experiment

This is a note to invite you–and any of your colleagues and students who may be interested–to participate in an interesting SC study. It is being run by Lonni Besançon and Jouni Helske. It’s an online survey that asks questions about CIs and other displays. It’s easy, takes around 15 minutes, and, as usual, is anonymous. To start, click here.

Feel free to pass this invitation on to anyone who might be interested. I suggest that we should all feel some obligation to encourage participation in SC research, because it has the potential to enhance research. Here’s to evidence-based practice! With cognitive evidence front and central.


Statistical Cognition, Some Background

Ruth Marom, Fiona Fidler, and I wrote about SC some years back. The full paper is here. The citation is:

Beyth-Marom, R., Fidler, F., & Cumming, G. (2008). Statistical cognition: Towards evidence-based practice in statistics and statistics education. Statistics Education Research Journal, 7, 20-39.

Some of Our SC Research

Here are a few examples:

Lai, J., Fidler, F., & Cumming, G. (2012). Subjective p intervals: Researchers underestimate the variability of p values over replication. Methodology: European Journal of Research Methods for the Behavioral and Social Sciences, 8, 51-62. Abstract is here.

Coulson, M., Healey, M., Fidler, F., & Cumming, G. (2010). Confidence intervals permit, but do not guarantee, better inference than statistical significance testing. Frontiers in Quantitative Psychology and Measurement, 1:26. Full paper is here.

Belia, S., Fidler, F., Williams, J., & Cumming, G. (2005). Researchers misunderstand confidence intervals and standard error bars. Psychological Methods, 10, 389-396. Abstract is here.

I confess that all those studies date from pre-Open-Science times, so there was no preregistration, and little or no replication. Opportunity!

Geoff

Play, Wonder, Empathy – Latest Educational Trends, Says The Open University

My long-time friend and colleague Mike Sharples told me about the recently released Innovating Pedagogy 2019 report from The Open University (U.K.). It’s the seventh in an annual series initiated by Mike. Each report aims to describe a number of promising trends in learning and teaching. There’s not much by way of formal evaluation of effectiveness and outcomes, but there are illuminating examples, and leads and links to resources to help adoption and further development.

The 2019 report describes 10 trends, as listed below. At this website you can click for brief summaries of any of the 10 that takes your fancy. There are also links to the previous six reports.

It strikes me that several of the 10 deserve thought, from the point of view of improving how we teach intro statistics. The one that immediately caught my eye was wonder.

I’ve always found randomness, and random sampling variability, to be the source of wonder. People typically don’t appreciate the wonder of randomness, nor do they appreciate that, in the short term, randomness is totally unpredictable and often surprising, even astonishing. In the long term, however, the laws of probability dictate that the observed proportions of particular outcomes will be very close to what we expect.

Prompted by the examples and brief discussion in the report of wonder, I can think of my years of work with the dances (of the means, of the CIs, of the p values, and more) as aiming to bring the wonder of randomness to students. Often we’ve discussed patterns and predictions and the hopelessness of making short-term predictions. We’ve compared the dances we see on screen–dancing before our eyes–with physical processes in the world that we might regard as random. (To see the dances, use ESCI, or go to YouTube and search for ‘dance of the p values’ and ‘significance roulette’.)

I suggest it’s worth poking about in this latest report, and in the earlier reports, for trends that might spark your own thinking about statistics teaching and learning.

Geoff

The ten 2019 trends:

Playful learningEvoke creativity, imagination and happiness 

Learning with robotsUse software assistants and robots as partners for conversation

Decolonising learningRecognize, understand, and challenge the ways in which our world is shaped by colonialism

Drone-based learningDevelop new skills, including planning routes and interpreting visual clues in the landscape

Learning through wonderSpark curiosity, investigation, and discovery

Action learningTeam-based professional development that addresses real and immediate problems

Virtual studiosHubs of activity where learners develop creative processes together 

Place-based learningLook for learning opportunities within a local community and using the natural environment

Making thinking visibleHelp students visualize their thinking and progress

Roots of empathyDevelop children’s social and emotional understanding

Open Science DownUnder: Simine Comes to Town

A week or two ago Simine Vazire was in town. Fiona Fidler organised a great Open Science jamboree to celebrate. The program is here and a few of the sets of slides are here.

Simine on the credibility revolution

First up was Simine, speaking to the title THE CREDIBILITY REVOLUTION IN PSYCHOLOGICAL SCIENCE. Her slides are here. She reminded us of the basics then explained the problems very well. Enjoy her pithy quotes and spot-on graphics.

My main issue with her talk, as I said at the time, was the p value and NHST framework that she used. I’d love to see the parallel presentation of the problems and OS solutions, all set out in terms of estimation. Of course it’s easy to cherry-pick and do other naughty things when using CIs, but, as we discuss in ITNS, there should be less pressure to p-hack, and the lengths of the CIs give additional insight into what’s going on. Switching to estimation doesn’t solve all problems, but should be a massive step forward.

A vast breadth of disciplines

Kristian Camilleri described the last few decades of progress in history and philosophy of science. Happily, there’s now much HPS interest in the practices of human scientists. So there’s lots of overlap with the concerns of all of us interested in developing OS practices.

Then came speakers from psychology (naturally), but also evolutionary biology, law, statistics, ecology, oncology, and more. I mentioned the diversity of audiences I’ve been invited to address this year on statistics and OS issues–from Antarctic research scientists to cardiothoracic surgeons.

Mainly we noted the commonality of problems of research credibility across disciplines. To some extent core OS offers solutions; to some extent situation-specific variations are needed. A good understanding of the problems (selective publication, lack of replication, misleading statistics, lack of transparency…) is vital, in any discipline.

IMeRG

Fiona’s own research group at The University of Melbourne is IMeRG (Interdisciplinary MetaResearch Group). It is, as its title asserts, strongly interdisciplinary in focus. Researchers and students in the group outlined their current research progress. See the IMeRG site for topics and contact info.

Predicting the outcome of replications

Bob may be the world champion at selecting articles that won’t replicate: I’m not sure of the latest count, but I believe only 1 or 2 of the dozen or so articles that he and his students have very carefully replicated have withstood the challenge. Only 1 or 2 of their replications have found effects of anything like the original effect sizes. Most have found effect sizes close to zero. 

Several projects have attempted to predict the outcome of replications, then assessed the accuracy of the predictions. Fiona is becoming increasingly interested in such research, and ran a Replication Prediction Workshop as part of the jamboree. I couldn’t stay for that, but she introduced it as practice for larger prediction projects she has planned.

You may know that cyberspace has been abuzz this last week or so with the findings of Many Labs 2, a giant replication project in psychology. Predictions of replication outcomes were collected in advance: Many were quite accurate. A summary of the prediction results is here, along with links to earlier studies of replication prediction.

It would be great to know what characteristics of a study are the best predictors of successful replication. Short CIs and large effects no doubt help. What else? Let’s hope research on prediction helps guide development of OS practices that can increase the trustworthiness of research.

Geoff

P.S. The Australasian Meta-Research and Open Science Meeting 2019 will be held at The University of Melbourne, Nov 7-8 2019.

Cabbage? Open Science and cardiothoracic surgery

“The best thing about being a statistician is that you get to play in everyone’s backyard.” –a well-known quote from John Tukey.

Cabbage? That’s CABG–see below.

A week or so ago Lindy and I spent a very enjoyable 5 days of sun, surf, and sand at Noosa Heads in Queensland. I spoke at the Statistics Day of the Annual Scientific Meeting of ANZSCTS (Australian and New Zealand Society of Cardiothoracic Surgeons). The program is here (scroll down to p. 18).

My first talk, to open the day, was “Setting the scene–problems with current design, analysis and reporting of medical research”. The slides are here.

In the afternoon I spoke on “‘Open science’–the answer to the problem?”. The slides are here.

Once again, I learned that:

  • The problems of selective publication, lack of reproducibility, and lack of full access to data and materials are, largely, common across numerous disciplines. And many researchers have increasing awareness of such problems.
  • Familiar Open Science practices (preregistration, open materials and data, publishing whatever the results, …) have wide applicability. However, each discipline and research field needs to develop its own best strategies for achieving, as well as it can, Open Science goals.

Technology races on…

I referred to a 2018 meta-analysis (pic below) that combined the results of 7 RCTs that compared two ways to rejoin the two halves of the sternum (breast bone) after open-chest surgery. The conclusion was that there’s not much to choose between wires and traditional suturing.

That was a 2018 article, but two commercial exhibitors were touting the advantages of devices that they claimed were better than either procedure assessed in the Pinotti et al. review. One was a metal clamp that has, apparently, been used for thousands of patients in China and has just been approved for use in Australia, on the basis of one RCT. The second looked like up-market plastic cable ties.

Open Science may set out ideal practice for researchers, but meanwhile regulators and practitioners must constantly make judgments on the basis of possibly less than ideal amounts of evidence, less than desirable levels of precision of estimates.

PCI or CABG? Just run a replication!

PCI is percutaneous coronary intervention, usually the insertion of a stent in a diseased section of coronary artery. The stent is typically inserted via a major blood vessel, for example the femoral artery from the groin.

CABG (“Cabbage”) is the much more invasive coronary artery bypass grafting, which requires open-chest surgery.

How do they compare? Arie Pieter Kappetein told us the fascinating story of  research on that question. He described the SYNTAX study, a massive comparison of PCI and CABG that involved 85 centres across the U.S. and Europe. At the 5-year follow-up stage, little overall difference was found between the two very different techniques. Some clinical advice could be given. There were many valuable subgroup analyses, some of which gave only tentative conclusions.

Replication was needed! More than 5 years and $80M later, he could describe results from the even larger EXCEL study. Again, there were many valuable insights and little overall difference, and the researchers are now seeking funding to follow the patients beyond 5 years. Recently his team has published a patient-level meta-analysis of results from 11 randomised trials involving 11,518 patients. Some valuable differences were identified and recommendations for clinical practice were made but, again, there was little overall difference in several of the most important outcomes–such as death.

So, in some fields, replication, if possible at all, is rather more challenging than simply running another hundred or so participants on your simple cognitive task!

Databases

Some of the most interesting papers I attended were retrospective studies of cases sourced from large patient databases. Such databases, as large and detailed as possible, are a highly valuable research resource. One seminar was devoted to the practicalities of setting up a major thoracic database, alongside the existing Australian cardiac database. The vast range of practicalities to be considered made clear how challenging it is to set up and keep running such databases.

Co-incidentally, The New Yorker that week published a wonderful article by Atul Gawande–one of my favourite writers–with the title Why Doctors Hate Their Computers. It seemed to me so relevant to that day’s cardiothoracic database discussions.

I hope you never have to worry about whether to prefer PCI or cabbage!

Geoff

A Wonderful Panorama of Statistics

Bob and I have been off-air for a while, but we haven’t gone away. I’ve been meaning for ages to blog about a wonderful book. Here it is:

Sowey, E., & Petocz, P. (2017). A panorama of statistics: Perspectives, puzzles and paradoxes in statistics. Wiley.

And the flyer with succinct information about the book is here. (Enjoy the full spread of the fine artwork that wraps around the book’s cover. The original painting, by Jeffrey Smart, is an enormous and wonderful sight in the foyer of one of Melbourne’s prominent theatres. Worth visiting!)

Panorama has been my beside-the-bed book for a while now. You could read it straight through, but I’ve preferred to dip in haphazardly, just about always finding something intriguing. It’s a cornucopia of statistical ideas, examples, oddities, paradoxes, historical tales, and more.

The back story: The journal Teaching Statistics, from 2003 to 2015 published the Statistical Diversions column by Peter Petocz and Eric Sowey. Peter and Eric are distinguished statisticians–and statistics teachers–based in Sydney.

Whenever a new issue of Teaching Statistics arrived I would first turn to their column to check out the new goodies, and the commentaries they gave on the questions they’d posed in the previous issue.

Eric and Peter now present the content of those columns, and more, assembled into coherent chapters as their Panorama book. It’s a great resource for any teacher looking for ways to engage or extend their students, or for anyone simply interested to explore–and be fascinated by–the discipline of statistics.

Here are a few tastes:

Randomness
Over about 20 years I built ESCI and wrote two books (the second with Bob, of course). For all that time I played with simulations of randomness, notably ESCI’s dances–of CIs in particular. I concluded early on that randomness is endlessly surprising and fascinating. It’s amazingly lumpy in the short term, while in the very long term fits exactly with what theory says we should expect. Even with this long experience, I found Chapters 11 (Some paradoxes of randomness) and 12 (Hidden risks for gamblers) especially interesting.

My brief version of Q11.5 (p. 89): Two people, Alice and Bert, toss a fair coin numerous times. Alice scores a point when a Head turns up, Bert a point for a Tail. How often is the lead likely to change? See pp.243-244 for the authors’ discussion–which may help us avoid unwarranted conclusions about what the movements of stock prices mean.

Getting the answer you want
You are teaching about questionnaires and wish to explain how a sequence of slanted questions can steer respondents in any direction you choose. A short and sharp satirical example is from the classic British Yes Minister program. See pp. 64 and 228 in the book, and the video here. (There are numerous links in the book. A list of all those links, in clickable form, is here.)

Eponymy and Stigler’s Law
We’re all familiar with many statistical eponyms (the Fisher exact test…). What is the relevance of Stigler’s law? Is that law true or false? If true it must be misleadingly named? See pp. 178-181 for an intriguing discussion, and pp. 292-295 for discussion of the questions posed in the earlier pages.

For more about the book, see an interview with the authors here, and to see the first few dozen pages of the book go here and click ‘look inside’.

Enjoy!
Geoff

P.S. On a totally different topic, one of the reasons I’ve been off-air is that Lindy and I joined a two-week tour of Greenland. It was fascinating. For example we visited the Ilulissat Glacier, which drains about 7% of the huge Greenland icecap, and which may have been the source of the iceberg that sank the Titanic. The most scary statistics I’ve seen for some time describe how that giant glacier–and others in Greenland–have greatly increased their rate of retreat in the last 10-20 years. In the case of the Ilulissat Glacier the calving is now no longer from a vast floating ice tongue, but from the much-retreated glacier front sitting on solid rock. So now the massive new icebergs all contribute to rising sea levels. That’s accelerating climate change in action, which is truly scary.

It’s not just Psychology: Questionable Research Practices in Ecology

Today’s fine article from The Conversation is:

Our survey found ‘questionable research practices’ by ecologists and biologists – here’s what that means

The authors are Fiona Fidler and Hannah Fraser, of The University of Melbourne.

Fidler and Fraser surveyed 807 researchers (494 ecologists and 313 evolutionary biologists) about their use of Questionable Research Practices (QRPs), including cherry picking statistically significant results, p-hacking, and hypothesising after the results are known (HARKing). The authors also asked them to estimate the proportion of their colleagues that use each of these QRPs. For each QRP, roughly around half the respondents stated that they had used that practice at least once. For some practices, they estimated higher rates among their research colleagues. These results are confronting, but the proportions are similar to those previously reported for psychology.

The preprint that gives more details of their survey and the results is here.

So QRPs have been endemic in Psychology, and now Ecology and Evolutionary Biology. And in even more disciplines, we’d have to guess. Open Science has, of course, developed to improve research practices, in particular by reducing QRPs markedly.

One of the problems is that anti-science forces can attempt to exploit these sort of findings, not to mention the also confronting findings of the replication crisis. The specific focus of Fidler and Fraser’s article is to respond to this problem. They pose and then reply to a number of the accusations that might be prompted by their results:

It’s fraud!
NO, it’s not! Scientific fraud does occur, and is extremely serious, but the evidence is that, thankfully, it’s very rare.

Scientists lack integrity and we shouldn’t trust them
The authors present evidence and several reasons why this is not true. The rapid rise and spread of Open Science may be the strongest indicator that researchers are responding with great integrity, energy, and conviction as they develop and adopt the better ways of Open Science.

We can’t base important decisions on current scientific evidence
On the contrary, in numerous important cases, including climate change and the effectiveness of vaccination, the evidence is multi-pronged, massive, and much replicated.

Scientists are human and we need safeguards
Yes indeed, and perhaps one of the biggest challenges of Open Science is to achieve change in the incentive systems that scientists are subjected to, and that so easily lead to QRPs.

But read the article itself–it’s short and very well-written.

Geoff

Randomistas: Dare we hope for evidence-based decisions in public life?

I’ve just listened to a great 20-min podcast, published by The Conversation. The podcast is here. It’s an interview by my colleague Fiona Fidler with Anthony Leigh, about his recently released book:

Randomistas: How Radical Researchers Changed Our World. Published by Black Inc. and La Trobe University Press.

Andrew Leigh is a Harvard-trained economist who was formerly a professor of economics at the Australian National University in Canberra. In Randomistas he argues that we should be using randomised trials much more often to guide public policy choices. He describes numerous examples of randomised trials, in a wide variety of fields. He’s well aware of the replication crisis and the Open Science practices needed to ensure trustworthy research.

So far, so good. But the really great thing is that Leigh is not just any ex-professor. He’s also an elected member of the House of Representatives, which is Australia’s Lower House of Parliament–approx. equiv. to Congress, or the House of Commons. Furthermore, he’s the Shadow Assistant Treasurer. If, as current polls suggest is likely, there is a change of government at the next Federal Election, due probably in early 2019, then he could easily be Australia’s Assistant Treasurer. And thus in a position to practise what he’s preaching in Randomistas.

Of course, it’s much easier to express good intentions when in Opposition than to put them in to practice when in Government. But it’s a great start that someone in his position knows enough, and cares enough about randomised trials and evidence-based policy-making to write so impressively about them.

Australia has had more than its share of atrocious political decisions that fly in the face of science and evidence. Dare we hope that a change of government might lead to an improvement?

Geoff