How we grade evidence

Every supplement on supplist gets graded on the things it's actually been studied for. Magnesium isn't graded as one thing — it's graded separately for blood pressure, sleep, constipation, muscle cramps, and so on. Each supplement is graded by health effect, and each health effect gets its own letter grade based on the human research.

Two tracks: efficacy and safety

On supplement pages, benefit claims and safety notes live on two separate tracks.

Efficacy is graded A through D, with when there isn't enough human research yet. This answers: does the human evidence support a benefit for this health effect?

Safety uses a separate track: ✓ SAFE, ⚠ CONCERNING, or ? UNKNOWN. Safety signals are evaluated separately from benefits, so no supplement gets one single overall score.

The efficacy scale

A
A · Strong evidence of benefit
B
B · Likely helpful
C
C · Mixed or limited evidence
D
D · Does not appear to help
— · Not enough human evidence yet

How grades are determined

Each health effect is graded on two axes: certainty of the evidence and the effect the evidence shows. The key idea is simple: each supplement is graded by health effect, not given one overall score.

Certainty

How sure are we about what the evidence says? Cochrane reviews and high-quality systematic reviews with formal risk-of-bias assessment can support moderate certainty. Direct RCT evidence without review-level synthesis sits at low. Mechanistic or animal-only evidence is very low. High certainty is rare from abstract-only data.

Effect

What does the best evidence show? Classified as meaningful positive, small positive, mixed, null, or negative. The classification comes from the single highest-quality review — Cochrane first, then top-tier systematic reviews. We don't vote-count across heterogeneous trials.

The two axes combine into a final grade for that health effect. Moderate certainty plus meaningful positive yields A. Low certainty plus mixed evidence yields C. Moderate certainty plus consistently null trials yields D.

An older percentage-style approach is not what these pages use now. The current methodology uses a certainty × effect matrix anchored to the strongest relevant human evidence.

Evidence hierarchy

When multiple types of evidence exist for a health effect, we anchor the grade to the strongest:

  1. Cochrane reviewsThe gold standard. Cochrane reviews use standardized risk-of-bias tools and structured certainty ratings. When a relevant Cochrane review exists, it carries the most weight.
  2. Systematic reviews and meta-analysesWith formal risk-of-bias assessment.
  3. Direct RCT bodyWhen no review-level evidence exists, we work from the trials directly.
  4. Observational studiesUsed to support but rarely to drive an efficacy grade.
  5. Mechanistic / animal evidenceInformative for plausibility, but not enough to establish efficacy in humans.

A Cochrane review on the wrong population (different dose, different age group, different health effect measure) doesn't get full weight. We track scope match and downgrade reviews whose scope doesn't match the claim being graded.

The abstract-only constraint

We use structured data built from PubMed abstracts. This lets us cover hundreds of papers per health effect, but abstracts don't include everything — full risk-of-bias details, exact effect sizes, allocation methods, and so on usually live in the full paper.

Practical implication: from abstracts alone, "high certainty" is rare. Most health effects top out at moderate certainty even when the underlying research is strong. The grading scale is built around this constraint.

Editorial overrides

The algorithmic grade is the starting point, not the final answer. A small number of grades are manually overridden when the algorithm's output clearly doesn't match the actual evidence — for example, when a health effect category is broader than the supporting reviews, or when a certainty rating is misclassified.

Every override is logged with the original algorithmic grade, the override grade, the reason, and the date.

The safety scale

✓ SAFE
No clear harm signal in the available human evidence.
⚠ CONCERNING
The human evidence points to a meaningful reason for caution.
? UNKNOWN
Not enough safety evidence to make a confident statement.

Not every label appears on the current supplement pages, but these are the safety labels the methodology supports.

What we don't do

We don't sell supplements ourselves. We don't run sponsored content. We don't accept payment from supplement brands to influence grades — the grade comes from the research, full stop.

Once a supplement clears our evidence bar, we may link to specific products from third-party retailers and earn affiliate commissions if you buy through those links. Affiliate relationships never affect which supplements get graded or what grade they receive. Products are recommended based on independent quality criteria — third-party testing, formulation, dose accuracy — not on which retailer pays the highest commission. See our affiliate disclosure for details.

Every claim traces back to peer-reviewed research indexed in PubMed.

We don't grade what hasn't been studied in humans. Mechanistic plausibility, animal data, and theoretical effects don't earn letter grades.

We don't promise effects in individuals. Letter grades describe what the evidence says about populations on average. Your response may differ.

Honest limitations

Letter grades are a simplification. Real evidence comes with confidence intervals, dose dependencies, population variations, and form-specific quirks. We try to surface the most important caveats on each supplement page, but a single letter is always less informative than reading the underlying research.

Evidence evolves. Every grade reflects the literature as of our last update. New trials and reviews may shift grades up or down — typically once or twice a year.

We're not a substitute for medical advice. Talk to your doctor before starting any supplement, especially if you take prescription medications, are pregnant, or have a chronic condition.