Read the Evidence. Not Just the Headline.
A practical guide to p-values, raw change, effect size, Hedges’ g, study design, and responsible application.
What p < .001 Actually Means.
If there were truly no between-group difference under the statistical model, a result at least this extreme would be expected less than about once in 1,000 from random variation alone.
How compatible are the data with “no true difference”?
A smaller p-value generally indicates the observed data are less compatible with that no-difference model, assuming the analysis and study conditions are appropriate.
How big, how useful, or how universal?
It does not measure magnitude, prove causation by itself, guarantee replication, eliminate bias, or predict an individual athlete.
Significance Tests the No-Difference Model. Effect Size Describes the Separation.
A statistically significant result can still be too small to matter in practice. Raw change and standardized magnitude restore the practical scale.
A Standardized Mean Difference with a Small-Sample Correction.
Hedges’ g places a group difference on a standardized scale and applies a correction intended to reduce upward bias in smaller samples.
It is unitless.
That makes it possible to discuss strength, repetitions, and throw distance on a common standardized scale.
It is not “times better.”
g = 3.99 does not mean 3.99× more effective. It describes standardized group separation.
It needs the raw change.
Always pair g with kilograms, pounds, repetitions, or meters so the reader can see what actually changed.
It needs context.
Sample size, variability, protocol, outcome, and population all shape interpretation.
Acute Mechanics and Training Adaptation Are Different Questions.
What changes during the rep?
Strong for comparing conditions within the same lifter. It does not establish long-term strength, hypertrophy, or injury outcomes.
What accumulates over weeks?
Stronger for a defined adaptation under a defined program. It does not automatically reveal the mechanism or generalize beyond the studied population.
Six Common Research Mistakes—Corrected.
Accurate
p < .001 describes how unusual a result at least this extreme would be under a no-difference model and its assumptions.
Accurate
Hedges’ g is a standardized mean-difference. It is not a percentage, multiplier, or probability.
Accurate
acute sEMG describes muscle-activation amplitude under that test. Hypertrophy requires longitudinal measurement.
Accurate
in a controlled training trial, the key question is whether the average improvement differed between groups under matched programs.
Accurate
it means the study did not demonstrate statistical significance for that outcome under that design and sample. Precision and power matter.
Accurate
randomization reduces allocation bias but does not erase sample limitations, measurement error, attrition, or context.
Ten Questions Before You Repeat the Headline.
Who was studied?
Sex, age, training status, sport, and sample size.
What was compared?
The exact intervention and the exact control condition.
What was randomized?
Participants, order, or condition—and how allocation was concealed.
What stayed the same?
Program, load, volume, coaching, testing, and recovery instructions.
What was measured?
The primary outcome and whether it matches the public claim.
What changed in raw units?
Kilograms, pounds, repetitions, meters, m/s, %MVC, or centimeters.
What did the p-value say?
The strength of evidence against a no-difference model—not magnitude.
What did the effect size say?
The standardized magnitude—not a multiplier.
What was not measured?
Pain, injury, hypertrophy, joint force, or real-world outcomes may remain unknown.
Does it fit my decision?
Transfer the finding only as far as the population and protocol justify.