What Pierre says, from Paris

pierre-d-dice-multixp-teaser2-14-nov-16Pierre Dragicevic (that’s his pic of a scary die!) is a super-interesting and enthusiastic researcher in HCI (Human-Computer Interaction) based at the Université Paris-Sud, an hour or so south of Paris. He is a researcher in the AVIZ Visual Analytics Project. He hosts a great page headed Bad Stats: Not What It Seems.

His page is definitely worth a browse! Here are some particular goodies:

Not one dance, but a dozen! Pierre has made a wonderful panel of a dozen dances that illustrates how successive results bounce around, simply because of sampling variability. Means, CIs, cat’s eyes, p values (bottom left), and much more. The key point, of course, is that any single CI (or, better still, cat’s eye) gives good information about the amount of dancing, whereas a single p value gives us virtually no information about the amount of sampling variability.

Scroll down a little to see details of his keynote talk at BioVis 2016. You can download his PowerPoint slides, with nice animations and a strong new-statistics message.

A little further down the page is a link to his chapter that discusses good statistical practices in the applied research field of HCI.

I’ll finish with another of Pierre’s telling pictures. This one references the front cover of the most famous book on usability, Don Norman’s The Design of Everyday Things. Yes, estimation is way more usable for the researcher than trying to pour coffee from the NHST pot!

Geoff pierre-d-multixp-teaser_3-14-nov-16

Geoff on Sydney radio: p hacking

Wendy Harmer is a great comedian and also a high-rating host on Sydney chat radio. Recently I (Geoff) had great fun chatting with her about p hacking, following my article in The Conversation on that topic. (See earlier post.)

Here’s our chat:

 

Trying to keep it simple, I describe just 3 tools, or questions, that are worth having in the toolkit for the intelligent and sceptical reader of a report of a research result, in the mass media. They are the take-home message. Keep handy at all times!

Geoff

Are you hearing ‘The Conversation’? This time on p hacking.

‘The Conversation’ is a non-profit online publication with the tag line ‘academic rigour, journalistic flair’. Starting in Australia a few years back and, at first, largely funded by universities, it is spreading worldwide, with editions in a number of countries, and recently a global edition. Articles are mainly written by academics (university faculty) with much editing for accessibility. It’s all free, and I find the daily email bulletins increasingly timely and interesting. It covers an incredible range of topics. Follow the US Presidential  campaigns, Brexit, climate change, or what you will. Timely and intelligent.

I mention this as a preview to mentioning my post of a couple of hours ago on p hacking and the replication crisis. I find an excuse to mention estimation and the new statistics. But you expected that. Feel free to add a comment in response to my post if you wish.

Sample sizes are too dang small…

Here’s another incredible paper by John Ioannidis and associates.  This one uses text mining to examining the statistical results of thousands of cognitive neuroscience and psychology papers.  It finds that sample sizes being used remain far too small: the typical paper has power of 0.12, 0.44, and 0.73 to detect small, medium, and large effect sizes, respectively.   In addition, there seem to be lots of statistical errors: 14% of papers report a result as statistically significant even if the underlying stats do not reach statistical significance!  Based on this analysis, the authors conclude that it really is likely that more than half of “significant” findings are false positives.  Depressing…but also a call to action to abandon NHST, embrace the New Statistics and Open Science, and to do better…we really can do better!

http://biorxiv.org/content/early/2016/08/25/071530

 

More on the dangers of p values

Here is an interesting new paper in which experienced researchers were asked to make judgements about research results expressed using the NHST approach (p values!).  A free copy of the paper is here, but it is easier to read the summary the authors wrote up here.

Here is how the authors summarize their findings:

“[1] Researchers interpret p-values dichotomously (i.e., focus only on whether p is below or above 0.05).
[2] They fixate on them even when they are irrelevant (e.g., when asked about descriptive statistics).
[3] These findings apply to likelihood judgments about what might happen to future subjects as well as to choices made based on the data.
We also show they ignore the magnitudes of effect sizes.”