Research literacy for coaches, athletes, media, and facilities

Read the Evidence. Not Just the Headline.

A practical guide to raw outcomes, p-values, confidence intervals, Cohen’s d, Hedges’ g, study design, and responsible interpretation—written so the statistics remain useful outside a laboratory.

Module 01 · The reading order

One Result. Five Different Questions.

The cleanest way to read exercise-science results is to separate the questions instead of asking one statistic to answer all of them.

01 · Raw outcome

What physically changed?

Pounds, kilograms, repetitions, meters, centimeters, m/s, watts, or %MVC. Start here because this is the result an athlete can recognize.

02 · Precision

How uncertain is the estimate?

Confidence intervals and variability show how tightly or loosely the study estimated the effect.

03 · P-value

How compatible are the data with no difference?

The p-value is an evidence signal under a statistical model. It is not the size of the result.

04 · Effect size

How large was the standardized separation?

Cohen’s d and Hedges’ g express a difference relative to the variability used in the calculation.

05 · Design

What conclusion can the study support?

An acute crossover, a randomized training trial, and a meta-analysis answer different questions.

Plain-language version Ask: What changed? How sure are we? How big was it? Who was studied? What exactly was tested?

That sequence is more useful than stopping at “statistically significant.”

Module 02 · Statistical evidence and precision

What p < .001 Actually Means.

Under the study’s null model and assumptions, p < .001 means that data at least this extreme would be expected less than 0.1% of the time if there were no underlying difference of the kind tested. That is strong incompatibility with the null model—not a 99.9% probability that an intervention “works.”1

What it helps answer

How surprising are the data under “no difference”?

Smaller p-values generally indicate that the observed data are less compatible with the no-difference model, assuming the analysis and study conditions are appropriate.

What it does not answer

How large, useful, or universal was the result?

A p-value does not measure effect magnitude, prove causation by itself, eliminate bias, guarantee replication, or predict an individual athlete.

What adds precision

Read the confidence interval with the estimate.

A narrower interval generally indicates greater precision. A wider interval means a broader range of values remains compatible with the data and model.

In weight-room language p < .001 says the pattern is difficult to reconcile with “nothing different happened” under the model.

It still does not tell you whether the difference was 2 pounds or 20 pounds. That is why raw outcomes and effect sizes remain necessary.

Module 03 · Standardized effect sizes

Cohen’s d and Hedges’ g Speak a Similar Language.

Both express a difference in standard-deviation units. The exact calculation must match the study design, and Hedges’ g adds a correction intended to reduce small-sample bias.23

Cohen’s d

A difference expressed relative to variability.

At its simplest, Cohen’s d divides a difference between averages by a standard deviation. A d of 1.0 means the difference is about one standard deviation—not 100% improvement.

  • The formula changes for independent groups, paired observations, and crossover designs.
  • Because the denominator matters, d values calculated differently should not be ranked blindly.
  • The AMM acute crossover paper reports Cohen’s d values from 0.23 to 1.02.
Hedges’ g

The same basic idea with a small-sample correction.

Hedges’ g applies a correction to reduce the tendency of standardized mean differences to be biased upward in smaller samples. Hedges’ foundational work demonstrated that small-sample bias and derived an adjusted estimator.3

  • In larger samples, d and g are usually very similar.
  • In smaller samples, g is generally slightly smaller than the uncorrected d.
  • The AMM longitudinal trials report Hedges’ g values from 3.67 to 4.10.
Hedges’ g = standardized mean difference × small-sample correction It is unitless. It is not a percentage, probability, multiplier, or guarantee.
Module 04 · Practical magnitude

Statistical Significance and Effect Magnitude Are Not the Same Thing.

Cohen’s familiar 0.2, 0.5, and 0.8 guideposts are useful orientation points, but they are conventions—not universal biological laws. The outcome, population, design, and effect-size formula all influence interpretation.24

0.2Conventional small reference
0.5Conventional medium reference
0.8Conventional large reference
3.67–4.10AMM longitudinal values reported in the published trials

The visual gap is intentional, but it is not a linear “better-than” meter. A g of 4.0 does not mean four times the improvement produced by a g of 1.0.

Study 01 · Acute crossover

Cohen’s d = 0.23–1.02

Five outcomes ranged from moderate-to-large by conventional guideposts and were statistically significant. Concentric power was d = 0.23 with p = .071, so no statistically significant acute-power claim is made.

Read the acute study
Study 02 · Four-week RCT

Hedges’ g = 3.85

The Launch Pad group gained 18.4 kg on average versus 11.1 kg for the Standard Flat Bench group—a 7.3 kg / 16.1 lb between-group advantage under the matched protocol.

Read the four-week study
Study 03 · Eight-week RCT

Hedges’ g = 3.99, 3.67, and 4.10

The published paper reported exceptionally large standardized between-group effects across 1-RM strength, NFL-225 repetitions, and seated medicine-ball throw distance.

Read the eight-week study
Module 05 · Exercise-science context

How Unusual Is an Effect Size Near 4.0?

The literature supports calling the AMM longitudinal effect sizes exceptionally large. It does not support turning them into “four times better,” assigning an invented percentile, or claiming they are the largest effects ever reported.

The strongest defensible conclusion

Exceptionally large by conventional benchmarks and selected exercise-science comparisons.

A 2024 meta-analysis covering 295 strength-and-conditioning studies, 535 groups, and 6,710 participants reported a pooled intervention-only standardized mean difference of 0.55. After outlier screening, retained effects ranged from −0.83 to 4.7. Values near 4 therefore sit close to the extreme upper end of that broad reported range—not near the pooled center.6

Important comparison guardrail

The 2024 database used intervention-only standardized changes, while the AMM longitudinal papers report between-group Hedges’ g. Other syntheses below use still other populations, outcomes, and estimands. The table provides scale and context—not a head-to-head ranking.

Selected peer-reviewed exercise-science literature used to contextualize standardized effect sizes
Source Evidence set Reported standardized effect Why it matters here
Swinton et al., 2024 295 resistance-only or resistance-dominant studies; 535 groups; 6,710 participants. Pooled intervention-only SMD = 0.55; retained range −0.83 to 4.7 after outlier screening. Provides a large strength-and-conditioning database. Values near 4 are close to the upper extreme of the retained range.
Hagstrom et al., 2020 24 studies of resistance-training adaptations in women. Upper-body strength g = 1.70; lower-body strength g = 1.40. Trim-and-fill estimates: 1.07 and 0.52. Shows that pooled training effects can exceed 0.8 while remaining materially below the AMM longitudinal values.
Martínez-García et al., 2021 16 controlled resistance-training studies; 424 overhead athletes. Throwing-velocity ES = 1.10; 95% CI 0.64–1.57. Offers a sport-performance transfer comparison based on controlled trials.
Makaruk et al., 2024 35 studies; 777 elite athletes. Sport-specific performance SMD = 1.16; 95% CI 0.65–1.66. Provides a pooled effect from elite-athlete resistance-training literature.
Hammert et al., 2026 High-load versus low-load isotonic training for non-specific strength. Cohen’s d = 0.322; 95% CI −0.08 to 0.72. Illustrates that pooled differences between two active resistance-training approaches are often modest and uncertain.

Selected examples were chosen for methodological relevance and proximity to resistance training or sport performance—not because they produce a preferred marketing comparison. The calculations are not interchangeable.

Why universal labels are limited

“Large” depends on the outcome.

In 114 exercise-therapy studies for tendinopathy, context-specific modeled quartiles across all outcomes were approximately 0.34, 0.73, and 1.21. Thresholds varied substantially by outcome domain, reinforcing that effect sizes should be interpreted against relevant evidence rather than one universal scale.5

What “rare” can responsibly mean

Unusual enough to demand attention—and replication.

The current evidence supports “exceptionally large” and “near the upper extreme of a broad published resistance-training distribution.” It does not support an exact rarity percentage because the literature does not provide a directly comparable universal percentile for g = 3.67–4.10.

Module 06 · Precision, stability, and replication

A Very Large Effect Deserves Attention and Scrutiny.

An extreme standardized effect can reflect a large real separation, low variability, a tightly controlled protocol, a homogeneous sample, an unstable small-sample estimate—or some combination of those factors. Hedges’ correction addresses a known bias; it does not remove every source of uncertainty.

01

Keep the raw difference visible.

Effect size should never replace pounds, repetitions, meters, or the actual group means.

02

Inspect variability and confidence.

Large separation with narrow uncertainty is more stable than a large point estimate surrounded by wide uncertainty.

03

Read the sample and protocol.

Small, homogeneous, supervised studies may estimate a strong effect under conditions that do not transfer perfectly elsewhere.

04

Look for convergence and replication.

Three studies asking different questions are stronger than one isolated result, but independent replication remains valuable.

Responsible language

“Exceptionally large standardized effects under the tested protocols.”

Pairs magnitude with the protocol boundary and leaves the raw results attached.

Avoid

“Four times better,” “guaranteed,” or “the largest ever.”

Those phrases convert a standardized statistic into claims the statistic does not establish.

Module 07 · Study design

The Design Determines What the Result Can Mean.

Acute crossover

What changes during the rep?

Strong for comparing conditions within the same lifter. It does not establish long-term strength, hypertrophy, pain, or injury outcomes.

Randomized training trial

What accumulates over weeks?

Stronger for a defined adaptation under a defined program. It does not automatically reveal the mechanism or generalize beyond the studied population.

Systematic review / meta-analysis

What pattern appears across studies?

Useful for pooling evidence, but the result depends on the quality, comparability, estimand, and bias of the included studies.

Translate the claim

Eight Common Research Mistakes—Corrected.

Wrong: “p < .001 means a 99.9% chance it works.”

Accurate

p < .001 describes how unusual data this extreme would be under a no-difference model and its assumptions.

Wrong: “d = 0.8 means an 80% improvement.”

Accurate

Cohen’s d is expressed in standard-deviation units. It is not a percentage.

Wrong: “Hedges’ g = 4 means four times better.”

Accurate

Hedges’ g is a standardized mean difference. It is not a multiplier, probability, or performance ratio.

Wrong: “A very large effect means every athlete improved.”

Accurate

An effect size summarizes group separation. Individual responses can still vary.

Wrong: “More sEMG means more muscle growth.”

Accurate

Acute sEMG describes activation amplitude under the test. Hypertrophy requires longitudinal measurement.

Wrong: “A non-significant result proves no effect.”

Accurate

It means the study did not demonstrate statistical significance for that outcome under that design, sample, and analysis. Precision and power still matter.

Wrong: “Any two d or g values can be ranked directly.”

Accurate

Check the outcome, design, comparator, variance standardizer, and effect-size formula before comparing.

Wrong: “Randomized means perfect.”

Accurate

Randomization reduces allocation bias but does not erase measurement error, attrition, sample limitations, or context.

A coach’s reading order

Ten Questions Before You Repeat the Headline.

01

Who was studied?

Sex, age, training status, sport, health status, and sample size.

02

What was compared?

The exact intervention, comparator, equipment, and training conditions.

03

What was randomized?

Participants, order, or condition—and how allocation was handled.

04

What stayed the same?

Program, load, volume, coaching, testing, recovery instructions, and measurement methods.

05

What changed in raw units?

Kilograms, pounds, repetitions, meters, m/s, watts, %MVC, or centimeters.

06

How precise was the estimate?

Read confidence intervals, standard deviations, and sample size—not only the point estimate.

07

What did the p-value say?

The compatibility of the data with a no-difference model—not effect magnitude.

08

What did the effect size say?

The standardized magnitude, calculated for that specific design—not a multiplier.

09

What was not measured?

Pain, injury, hypertrophy, joint force, retention, or real-world outcomes may remain unknown.

10

Does it fit my decision?

Transfer the finding only as far as the population, protocol, and outcome justify.

References and source trail

Keep the Literature Attached to the Interpretation.

The statistical explanations and effect-size context on this page are grounded in methodological literature, exercise-science syntheses, and the three published AMM studies.

  1. Wasserstein, R. L., & Lazar, N. A. (2016). The ASA’s statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129–133. Open DOI
  2. Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology, 4, 863. Open DOI
  3. Hedges, L. V. (1981). Distribution theory for Glass’s estimator of effect size and related estimators. Journal of Educational Statistics, 6(2), 107–128. Open DOI
  4. Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates.
  5. Swinton, P. A., et al. (2023). What are small, medium and large effect sizes for exercise treatments of tendinopathy? A systematic review and meta-analysis. BMJ Open Sport & Exercise Medicine, 9(1), e001389. Open DOI
  6. Swinton, P. A., Schoenfeld, B. J., & Murphy, A. (2024). Dose–response modelling of resistance exercise across outcome domains in strength and conditioning: A meta-analysis. Sports Medicine, 54, 1579–1594. Open DOI
  7. Hagstrom, A. D., Marshall, P. W., Halaki, M., & Hackett, D. A. (2020). The effect of resistance training in women on dynamic strength and muscular hypertrophy: A systematic review with meta-analysis. Sports Medicine, 50(6), 1075–1093. Open DOI
  8. Martínez-García, D., et al. (2021). Strength training for throwing velocity enhancement in overhead throw: A systematic review and meta-analysis. International Journal of Sports Science & Coaching, 16(5). Open DOI
  9. Makaruk, H., Starzak, M., Tarkowski, P., Sadowski, J., & Winchester, J. (2024). The effects of resistance training on sport-specific performance of elite athletes: A systematic review with meta-analysis. Journal of Human Kinetics, 91, 135–155. Open DOI
  10. Hammert, W. B., et al. (2026). Non-specific strength changes between high- and low-load isotonic resistance training: A systematic review and meta-analysis. Sports Medicine, 56, 763–773. Open DOI
  11. Kidwell, J. A., et al. (2026). Acute effects of thoracic-spinal elevation via a novel bench press pad on sEMG and barbell kinetics in resistance-trained males. International Journal of Exercise Science, 19(1), 1003. Open DOI
  12. Goldman, P., et al. (2025). Eccentrically overloaded bench press training: Augmenting strength gains via a novel bench press pad. Scientific Journal of Sport and Performance, 4(4), 480–490. Open DOI
  13. Blatney, A. E., et al. (2026). Effects of an eight week training regimen with a novel bench press pad compared to a traditional bench on upper body strength and performance in collegiate American football players. Scientific Journal of Sport and Performance, 5(1), 10–21. Open DOI
Apply the framework

Now Read the Three Studies with the Context Attached.