Why Supplement Meta-Analyses Disagree

Nikos DrosakisFounder and responsible editor3 min read

A personal essay, not an evidence assessment. Our graded assessments of individual ingredients — written to a published standard, from full texts — are in the evidence section. Nothing here is a statement about what any product does.

If one study says yes and another says no, people often ask for the meta-analysis.

Understandably.

A meta-analysis sounds like the final court of appeal.

Many studies go in.

One answer comes out.

Unfortunately, it is not always that simple.

I have seen meta-analyses that appear to investigate the same supplement and reach noticeably different conclusions.

That does not necessarily mean one team made a mistake.

Often they are answering subtly different questions.

Which studies went in?

This is the first place to look.

One review may include:

healthy adults.

Another:

patients.

Another:

older adults.

Another may combine them.

One may include only randomised trials.

Another may include uncontrolled interventions.

One may exclude studies below a certain duration.

Another may not.

Before looking at the pooled number, I want to know what was pooled.

What counts as the same intervention?

With supplements, this can become difficult.

Different chemical forms.

Different extracts.

Different standardisation.

Different doses.

Different combinations.

Different durations.

If we put all of these into one bucket because they share an ingredient name, the resulting average can become biologically difficult to interpret.

Outcomes can be equally heterogeneous

One paper measures:

reaction time.

Another:

memory.

Another:

subjective fatigue.

Another:

a composite cognitive score.

If these are combined too aggressively under a heading like “cognitive function,”

precision can become cosmetic.

The decimal places become cleaner while the biological question becomes less clear.

Statistical heterogeneity is not something a model makes

disappear

A random-effects meta-analysis allows for heterogeneity.

It does not make heterogeneity cease to be a problem.

A statistical model cannot rescue a conceptually incoherent dataset.

Small studies can influence pooled conclusions

In random-effects models, smaller studies may receive relatively more weight than they would under a fixed-effect model.

If small-study effects or publication bias are present, this can influence the pooled result.

So when two meta-analyses choose different models or include different small trials, disagreement can emerge naturally.

Then there is publication bias

If positive small studies are easier to publish than negative small studies, the literature available for meta-analysis is already distorted.

The meta-analysis can be statistically sophisticated and still be analysing an incomplete reality.

Quality decisions matter

What does the review do with a poorly blinded study?

What about a trial with questionable randomisation?

What about a paper where outcomes do not match the registry?

Include?

Exclude?

Downgrade?

Run sensitivity analysis?

Reasonable researchers can make different defensible decisions.

I do not ask which meta-analysis is “the winner”

I ask:

Why do they disagree?

Sometimes that question teaches me more than either pooled result.

If the difference disappears when one unusual study is removed, that matters.

If one review includes clinical populations and another does not, that matters.

If one combines different formulations, that matters.

If heterogeneity is enormous, that matters.

A meta-analysis sits near the top of many evidence hierarchies for good reason.

But a meta-analysis is not alchemy.

Poor, incompatible or biased studies do not become excellent merely because we calculate a weighted average.

Sometimes the most responsible conclusion from a meta-analysis is not:

The answer is X.

It is:

The available studies are too different for X to mean what people want it to mean.

Methodological sources

Reporting and appraisal standards referred to in this essay. They are not the evidence behind any product claim.

Next in the seriesIndustry-Funded Research: When Should We Be Concerned?