Statistical Significance Is Not Clinical Significance

Nikos DrosakisFounder and responsible editor3 min read

A personal essay, not an evidence assessment. Our graded assessments of individual ingredients — written to a published standard, from full texts — are in the evidence section. Nothing here is a statement about what any product does.

There is one number that has probably sold more questionable scientific certainty than almost any other:

p < 0.05 It looks authoritative.

A threshold has been crossed.

The result is “significant.”

For someone reading quickly, the natural conclusion is:

It worked.

But statistical significance answers a much narrower question.

And it is very possible for a result to be statistically significant while being practically unimportant.

Imagine an enormous study

Suppose we measure reaction time in tens of thousands of people.

One group takes a supplement.

The other takes placebo.

The supplement group becomes, on average, fractionally faster.

Because the dataset is enormous, the analysis may detect that tiny difference with great statistical confidence.

The p-value can look impressive.

But the practical question remains:

Would anyone notice the difference?

If the answer is no, the statistical result may still be scientifically interesting.

It is not automatically a meaningful consumer benefit.

The opposite problem also exists

A study may observe a potentially meaningful difference and fail to reach conventional statistical significance.

Why?

Perhaps the sample is too small.

Perhaps variability is high.

Perhaps the confidence interval is wide.

That does not prove the effect exists.

But neither does it automatically prove there is no effect.

This is why I dislike binary interpretation:

significant = works non-significant = doesn't work Reality is usually less convenient.

I want to know the magnitude

If somebody tells me a result is statistically significant, my next question is:

How large is the effect?

Then:

What is the confidence interval?

The effect estimate tells us what happened in the sample.

The interval tells us something about the range of values compatible with the data.

That is considerably more informative than a ceremonial crossing of 0.05.

Multiple comparisons make things even more interesting

Imagine testing:

memory, attention, reaction time, working memory, mood,

fatigue, stress, accuracy, sleepiness, and several subscales.

The more statistical tests we perform, the greater the opportunity for something interesting to appear simply through chance.

This does not mean a significant secondary result is fake.

It means context matters.

That is also why prespecified primary outcomes matter so much.

“Significant” is an unfortunate word

In ordinary English, significant means:

important.

In statistics, it means something different.

That linguistic collision creates endless confusion.

A statistically significant result can be trivial.

A clinically or practically important possibility can remain statistically uncertain.

When I read papers, I therefore try to mentally replace:

statistically significant

with:

the data crossed this statistical decision threshold.

It is much less exciting.

And therefore much safer.

What do I want MindHeaven to report?

Ideally:

the observed difference, the effect size,

the confidence interval, the sample size, the primary outcome, the comparator, and the uncertainty.

Not merely:

p = 0.031.

A precise-looking number should never become a substitute for thinking.

The question is not “Was it significant?”

The question is:

How large was the effect, how certain are we that it is real, how robust is the study, and does the difference actually matter?

That is a more demanding question.

It is also the question the consumer thought we were answering in the first place.

Methodological sources

Reporting and appraisal standards referred to in this essay. They are not the evidence behind any product claim.

Next in the seriesWhat Does a 20-Person Study Actually Tell Us?