Research literacy for coaches, athletes, and media

Read the Evidence. Not Just the Headline.

A practical guide to p-values, raw change, effect size, Hedges’ g, study design, and responsible application.

Module 01 · Statistical significance

What p < .001 Actually Means.

If there were truly no between-group difference under the statistical model, a result at least this extreme would be expected less than about once in 1,000 from random variation alone.

What it helps answer

How compatible are the data with “no true difference”?

A smaller p-value generally indicates the observed data are less compatible with that no-difference model, assuming the analysis and study conditions are appropriate.

What it does not answer

How big, how useful, or how universal?

It does not measure magnitude, prove causation by itself, guarantee replication, eliminate bias, or predict an individual athlete.

Module 02 · Practical magnitude

Significance Tests the No-Difference Model. Effect Size Describes the Separation.

A statistically significant result can still be too small to matter in practice. Raw change and standardized magnitude restore the practical scale.

0.2Conventional small reference
0.5Conventional medium reference
0.8Conventional large reference
3.67–4.10AMM training-study values
Context-dependent reference points—not universal laws, not multipliers, and not individual guarantees.
Module 03 · Hedges’ g

A Standardized Mean Difference with a Small-Sample Correction.

Hedges’ g places a group difference on a standardized scale and applies a correction intended to reduce upward bias in smaller samples.

01

It is unitless.

That makes it possible to discuss strength, repetitions, and throw distance on a common standardized scale.

02

It is not “times better.”

g = 3.99 does not mean 3.99× more effective. It describes standardized group separation.

03

It needs the raw change.

Always pair g with kilograms, pounds, repetitions, or meters so the reader can see what actually changed.

04

It needs context.

Sample size, variability, protocol, outcome, and population all shape interpretation.

Module 04 · Study design

Acute Mechanics and Training Adaptation Are Different Questions.

Acute crossover

What changes during the rep?

Strong for comparing conditions within the same lifter. It does not establish long-term strength, hypertrophy, or injury outcomes.

Randomized training trial

What accumulates over weeks?

Stronger for a defined adaptation under a defined program. It does not automatically reveal the mechanism or generalize beyond the studied population.

Translate the claim

Six Common Research Mistakes—Corrected.

Wrong: “p < .001 means a 99.9% chance it works.”

Accurate

p < .001 describes how unusual a result at least this extreme would be under a no-difference model and its assumptions.

Wrong: “Hedges’ g = 4 means four times better.”

Accurate

Hedges’ g is a standardized mean-difference. It is not a percentage, multiplier, or probability.

Wrong: “More EMG means more muscle growth.”

Accurate

acute sEMG describes muscle-activation amplitude under that test. Hypertrophy requires longitudinal measurement.

Wrong: “Both groups improved, so the equipment did nothing.”

Accurate

in a controlled training trial, the key question is whether the average improvement differed between groups under matched programs.

Wrong: “A non-significant result proves no effect.”

Accurate

it means the study did not demonstrate statistical significance for that outcome under that design and sample. Precision and power matter.

Wrong: “Randomized means perfect.”

Accurate

randomization reduces allocation bias but does not erase sample limitations, measurement error, attrition, or context.

A coach’s reading order

Ten Questions Before You Repeat the Headline.

01

Who was studied?

Sex, age, training status, sport, and sample size.

02

What was compared?

The exact intervention and the exact control condition.

03

What was randomized?

Participants, order, or condition—and how allocation was concealed.

04

What stayed the same?

Program, load, volume, coaching, testing, and recovery instructions.

05

What was measured?

The primary outcome and whether it matches the public claim.

06

What changed in raw units?

Kilograms, pounds, repetitions, meters, m/s, %MVC, or centimeters.

07

What did the p-value say?

The strength of evidence against a no-difference model—not magnitude.

08

What did the effect size say?

The standardized magnitude—not a multiplier.

09

What was not measured?

Pain, injury, hypertrophy, joint force, or real-world outcomes may remain unknown.

10

Does it fit my decision?

Transfer the finding only as far as the population and protocol justify.

Apply the framework

Now Read the Three Studies with the Context Attached.