Library · Reading research · Explainer

How to read a training study

A training study is built out of about seven choices, and every one of them is written down in the paper. Once you can see them, you can tell how far a finding travels — no statistics required.

Published 18 August 2026 · rewritten 19 August 9 sources · 13 figures · every number checked against the source
The machinery underneath, in plain words
A stack of printed research pages on a pale wooden table seen from above, with a highlighter, reading glasses, a mug of coffee and the heel of a running shoe at the edge of the frame.
Figure 1. Illustration. Everything below is about this: a printed study, twenty minutes, and knowing where to look.

You cannot check from your kitchen table whether a study is true. You can work out how far it travels — and that is nearly always the more useful thing to know. Seven questions get you there, and all seven are answered in the paper itself.

  • Six of the seven sit in the methods section — the part most readers skip on the way to the conclusion. It is the most informative page in the paper.
  • Training studies are small — half of them in one flagship journal had nineteen people or fewer. That is normal, and it changes what you can fairly ask of a single paper.
  • "Compared with what" is the question that goes missing most often, because a headline has room for only half the sentence.
  • Around eight in ten published sport studies report a finding — which tells you something about how papers get published, and is worth knowing before you count them up.

Prefer to watch? All seven questions in 5:09, narrated, with subtitles you can switch off — it has its own page too. Every number in it is traced to a source below.

You read that runners improved three percent after eight weeks of something. The first question that comes to mind is usually is that true? — and that one is close to unanswerable from where you are sitting. There is a better question, and the paper will answer it for you: how was this measured?

A training study is assembled out of a small number of decisions. Who took part. How many of them. How long it ran. What the other group did instead. What got measured, and how big the difference turned out to be. Those decisions are what a finding is made of, and they decide how far it travels — whether it is about people like you, over a time span like yours, against the alternative you were actually considering.

They are also written down, nearly always in the same order, and none of them require any statistics to understand. What follows is where each one lives and what to make of it.

Diagram: a schematic research paper with its abstract, methods, results and discussion sections, and the seven questions mapped to where each is answered. Five sit in the methods section, one in the results, and the seventh — whether it would have been published anyway — is answered outside the paper entirely.
Figure 2. Where each answer lives. Five of the seven are in the methods section. The seventh is the only one the paper itself cannot tell you.

1 Who was in it?

The methods section opens by telling you who this finding is about, and that single paragraph decides most of what the study can mean for you.

The detail that moves the needle furthest is training status. Somebody who has never trained has a great deal of room to improve, and in that state almost any sensible programme produces a result. The closer someone is to what their body can do, the smaller the remaining gains and the harder they are to find. A study run on untrained volunteers is a perfectly good study; it is just answering a question about untrained volunteers.

Diagram: three columns showing current fitness against a dashed ceiling line marking what each body could eventually do. An untrained person has a large gap, a club runner a moderate one, an athlete near their ceiling almost none.
Figure 3. The same programme, three different amounts of room to work with. This is the single biggest reason a result found in beginners rarely arrives intact at the sharp end.

The rest of the paragraph is worth the same thirty seconds. Age, sex, and the level people were competing at all shape who the answer belongs to. None of this makes a study weaker — it makes it specific, which is what a study is supposed to be. The reading error is ours, when we quietly assume the participants were people like us.

A concrete one from our own shelf: the famous four-percent figure for carbon-plated racing shoes was measured on eighteen very fast men running at around two-and-a-half-hour marathon pace. That is a real result about those runners at that speed. Whether it is a result about a four-hour marathoner is a different question, and the paper never claimed to answer it. We took that apart in the super-shoes check.

2 How many took part?

Training studies are small. Not carelessly small — lab time is expensive, good athletes are scarce, and asking twenty people to change how they train for two months is already a serious undertaking.

One survey of the Journal of Sports Sciences put the median at nineteen participants (figure 4). A look at four other sport-science journals found averages between fifteen and thirty-two, with a spread wider than the average itself — which is another way of saying that the small studies are not outliers, they are the shape of the field.

Bar chart: median of 19 participants in the Journal of Sports Sciences, and mean sample sizes of 21.4, 15.0, 32.2 and 20.2 in four other sport-science journals, each with a standard-error bar. Under each bar the number of papers and the standard deviation: 31 papers and SD 23.6, 21 papers and SD 18.8, 29 papers and SD 31.8, 37 papers and SD 21.5.
Figure 4. How big a training study usually is. Whiskers are one standard error either way — that is our own calculation, SD ÷ √n, from the figures Knudson publishes. The median on the left has none, because a median does not have one. Worth noticing separately: the standard deviations printed under the bars are larger than the means themselves, so these groups vary enormously in size.

What matters is what that buys you. A small study is built to show a large effect. It is poorly equipped to rule out a small one. Give twenty athletes a medium-sized difference to find and the study lands on it fewer than half the times you run it; make the group forty-four and it lands four times in five (figure 5).

Chart: with 20 participants a study has a 45% chance of finding a real medium-sized effect; with 44 participants that rises to 80%. A dotted line marks the coin-toss level and a dashed line the 80% level researchers aim for.
Figure 5. The chance a study spots an effect that is genuinely there. Researchers call this statistical power; in plain terms it is how often the study would succeed if you ran it over and over.

That has one very practical consequence, and it is probably the most useful single sentence in this piece. When a small study reports no difference between the groups, that usually means there were too few people to see one — not that there is nothing there. A null result from twenty participants is a shrug, not an answer.

There is a design that makes small studies punch well above their size, and it is worth recognising: the crossover. Instead of comparing one group of people with another, every participant does both conditions and is compared with themselves. That removes the biggest source of noise in sport — the fact that people differ enormously from one another — and it is why a well-run crossover on fifteen runners can be more informative than two groups of thirty.

3 How long did it last?

Where a study stops decides what it is able to tell you, and the answer is usually one line in the methods: eight weeks, twelve weeks, one afternoon.

Two things ride on that number. The first is what could even have changed in the time available. Some adaptations move within a fortnight; others need months of consistent work before they show up at all. A six-week study that reports no change in something that takes half a year to shift has not found an absence — it stopped early.

The second is what happens after the last measurement, which is: nobody knows. Early gains in almost any new programme are partly the novelty of a new stimulus, and part of that fades. Whether what remains keeps building or settles back is a question the study did not ask.

Diagram: a curve rising over eight weeks, then splitting into two dashed continuations — one that keeps building and one that fades. The area past the final measurement is shaded and marked as unmeasured.
Figure 6. The solid line is what was measured; both dashed lines are consistent with it. Reasonable extrapolation is still extrapolation.

Methodologists take this seriously enough to have a rule about it: decide in advance which time points you care about, so that nobody chooses afterwards which measurement to report. When a paper measured at four, eight and twelve weeks and the headline quotes one of them, it is fair to wonder what the other two showed.

4 Compared with what?

Every result is a difference between two things, and the thing on the other side decides what the difference means. This is the question that disappears most reliably, because a headline has room for only half the sentence.

Methodologists split comparisons into two kinds. An inactive comparison is a placebo, a waiting list, no treatment at all, or simply carrying on as before. An active comparison is a real alternative — a different version of the same thing, or the cheaper option you would otherwise have chosen. Both are legitimate, and they answer different questions. Beating nothing tells you the thing does something. Beating the sensible alternative tells you whether to switch, and that is nearly always the question you actually have.

Diagram: four comparison types side by side — nothing at all, a dummy version, usual practice, and the real alternative — with a double-headed arrow beneath showing that the same result looks most impressive on the left and is most useful to you on the right.
Figure 7. The same difference, four different meanings. None of these is a flawed design; they simply answer different questions.

Sport adds a complication that medicine mostly avoids: you can rarely hide which group someone is in. You know whether you are wearing the stiff plated shoe or the flat one, and whether you got the drink or the water. Where the comparison is visible, part of the difference is expectation — and expectation moves performance by an amount you can measure. That is not a reason to set the study aside. It is a reason to read the comparison as carefully as the headline number.

The carbon-shoe research is a clean example again: the comparison there is usually a light racing flat, itself a fast shoe, rather than an ordinary trainer. The famous saving is a difference between two good shoes. Read against the wrong comparison, it sounds roughly twice as useful as it is.

Three things to look for. What did the other group actually do — nothing, or something plausible? Could either side tell which group they were in? And if the comparison is "usual training", whose usual training is that: yours, or that of the students who volunteered?

5 What was measured?

The outcome a study measures is rarely the outcome you have in mind, and the distance between the two is where a lot of headlines are made.

A runner on a treadmill in a sports-science laboratory, wearing a breathing mask connected by a corrugated tube to gas-analysis equipment, seen from the side.
Figure 8. Illustration. Most training research is measured in a room like this, which is exactly why the outcome on the page is usually a step or two away from a race result.

Take running economy — how much oxygen you burn to hold a given pace. It is a sensible thing to measure: it can be done in an afternoon, it is precise, and it is genuinely related to performance. But it is a stand-in. What you actually want to know is your time on race day, and that takes a season, a start line and a good deal of luck to observe.

Cochrane's methods handbook is blunt about stand-ins of this kind: laboratory results "may not predict clinically important outcomes accurately", and plenty of interventions improve the stand-in without improving the thing it stands in for. The same logic applies to a lactate value, a peak power number or a jump height. None of them are wrong. They are just one or two steps back from the finish line.

Diagram: a chain of four boxes running from oxygen cost at a set pace, to running economy, to the pace you can hold in a race, to your finish time. The first two are marked measured, the last two inferred, with a marker showing where the study stops.
Figure 9. Where the measuring stops and the reasoning starts. Everything to the right of the line may well be right — it was simply not observed.

One more thing hides in the same paragraph: which outcome the researchers said they would look at. A study that measures fifteen things will almost certainly find something interesting in one of them, and a finding that was chosen after the fact is a much weaker claim than one that was named in advance. Good papers say which was the primary outcome. It is worth checking whether the number in the headline is that one.

6 How big was the difference?

Here is the one place where a word does more damage than any other: significant. In ordinary English it means important. In a paper it means something much narrower — that a difference this size is unlikely to have come from chance alone.

Those two sentences say different things, and the gap runs both ways. A large study can find a difference so small that nobody would notice it, and call it significant with a straight face. A small study can miss a difference that would genuinely matter, and report no significant effect. Neither sentence tells you how big anything was.

For that, look for the effect size — a number that expresses how large the difference is relative to how much people vary anyway. Most papers report one, usually as a value with a letter in front of it, and the thresholds everybody quotes for small, medium and large came out of behavioural science decades ago. Sport now has its own: when one group pooled six and a half thousand effects from strength and conditioning trials, the same three words landed on noticeably smaller numbers (figure 10).

Bar chart comparing the textbook small, medium and large effect-size thresholds of 0.20, 0.50 and 0.80 with the sport-specific values of 0.16, 0.46 and 0.81 pooled from 643 strength and conditioning studies.
Figure 10. The same three words, two different scales. Using the field's own yardstick beats borrowing one from another discipline.

The best question of all, though, is more direct than any of this: how much smaller is the difference than the thing you would notice? Researchers call that threshold the smallest meaningful change — the point below which a difference, however real and however significant, is not worth reorganising your week for. Comparatively few sport papers name one, and noticing that it is missing is itself a useful reading skill.

The version of this you can always do yourself is to translate the percentage into time. A one-percent improvement in a forty-five-minute 10K is about twenty-seven seconds; in a four-hour marathon it is a little over two minutes. Whether that is worth anything is a question about your season, not about the study — but you cannot answer it while the result is still a percentage. (Our pace calculator will do the arithmetic.)

7 Would it have been published anyway?

This is the only one of the seven that the paper in front of you cannot answer, and it is the reason to hold every single study a little more loosely than it holds itself.

Research does not work out very often. Most careful ideas turn out to be wrong, or too small to see, or true only under conditions nobody can reproduce. So you would expect a fair share of published papers to report that they found nothing much. In sports and exercise medicine, a survey of 129 papers found that 82.2% reported a positive result; a wider check across kinesiology journals a year later landed on 81.4%, near enough the same number that the authors call the two indistinguishable (figure 11).

Bar chart: 82.2% of sports and exercise medicine papers report a significant finding, and 81.4% of papers across kinesiology journals — two separate surveys landing within a percentage point of each other.
Figure 11. Two teams, two sets of journals, the same answer. What you get to read is a selection of what was done.

That number is telling you about publishing rather than about training, and it is worth being precise about the mechanism, because it is not bad faith. It works from both ends. Journals prefer a paper with a result in it — it gets read, it gets cited, and a page reporting that nothing happened is a hard sell to an editor. And authors prefer to write up the study that found something, for the same reasons plus one more: the paper is easier to write when there is a story in it. Neither party is doing anything they would be ashamed of. The effect is still that the quiet results end up in a drawer, and the ones you read are the survivors.

This is called publication bias, and it is the reason a single striking paper deserves less weight than it feels like it deserves — you are seeing the studies that made it through the filter, not the ones that were run.

The genuinely good news is that the fix exists and you can watch it working. In a registered report, the question, the method and the analysis are reviewed before any data exist, and the journal commits to publishing whatever comes out. Nobody can file the dull result away, and nobody can reshape the hypothesis to fit the numbers. When one team compared 71 registered reports with 152 ordinary ones in the same journals, the share reporting a positive finding went from 96% to 44% (figure 12).

Bar chart: 96% of standard reports end up with a positive finding, versus 44% of registered reports where the plan is accepted before any data are collected.
Figure 12. Same researchers, same journals, different order of operations. Roughly half the positive findings turn out to depend on the plan being written after the data arrived.

A handful of sport journals now run the format, and when a paper you are reading is a registered report it says so at the top. That is a real mark of quality, and it costs you nothing to look for it.

8 The seven questions

None of these need statistics, and all of them except the last are answered somewhere in the paper. Together they tell you how far a finding travels, which is almost always more useful than knowing whether it cleared a threshold.

  1. Who was in it? Trained or untrained, and how much like you. This one moves the answer more than any other.
  2. How many took part? Small is normal. It means the study is built to show a large effect, and that "no difference" probably means "too few people to see one".
  3. How long did it last? Long enough for the thing being studied to change? And what happened after the last measurement — which nobody knows.
  4. Compared with what? Nothing, a dummy, usual practice, or the alternative you would actually have chosen. Only the last one answers your question.
  5. What was measured? And how many steps sit between that and the outcome you care about. Was it the outcome they said they would measure?
  6. How big was the difference? In seconds per kilometre, not in significance. Big enough that you would notice?
  7. Would it have been published anyway if the answer had been "no difference"? A registered report is the one design where you know the answer is yes.
The seven questions set out as a numbered checklist: who was in it, how many, how long did it last, compared with what, what was measured, how big was the difference, and would it have been published anyway.
Figure 13. Save this one. It fits on a phone screen and works on any training paper you are likely to meet.

A study is a photograph, not a promise. Read for who was in the picture, how many, how long the shutter was open — and what was standing just outside the frame.

Sources

This is an explainer rather than a review: the sources below were picked to explain how a study is put together, not gathered through a systematic search. Eight of the nine were read and checked at their own source rather than through somebody else's summary. That matters more than it sounds: doing it turned up three numbers on the first version of this page that needed correcting. Where a figure still reaches you second-hand, the entry says so.

  1. T1 Büttner F, Toomey E, McClean S, Roe M, Delahunt E. Are questionable research practices facilitating new discoveries in sport and exercise medicine? The proportion of supported hypotheses is implausibly high. Br J Sports Med. 2020;54:1365–1371. doi:10.1136/bjsports-2019-101863 — read at source; the 82.2% is checked against it.
  2. T1 Twomey R, Yingling V, Warne J, et al. The nature of our literature: a registered report on the positive result rate and reporting practices in kinesiology. Commun Kinesiol. 2021;1(3). doi:10.51224/cik.v1i3.43 — read at source; they report 81.43% [75.78–86.3].
  3. T1 Knudson DV. Authorship and sampling practice in selected biomechanics and sports science journals. Percept Mot Skills. 2011;112:838–844. doi:10.2466/17.PMS.112.3.838-844 — read at source; the four journal averages come from its table 1.
  4. T1 Abt G, Boreham C, Davison G, Jackson R, Nevill A, Wallace E, Williams M. Power, precision, and sample size estimation in sport and exercise science research. J Sports Sci. 2020;38:1933–1935. doi:10.1080/02640414.2020.1776002 — read at source; the median sample size of 19 is checked against it.
  5. T1 Scheel AM, Schijen MRMJ, Lakens D. An excess of positive results: comparing the standard psychology literature with registered reports. Adv Methods Pract Psychol Sci. 2021;4(2). doi:10.1177/25152459211007467 — read at source; 96% versus 44% checked against it.
  6. T1 Swinton PA, Burgess K, Hall A, Greig L, Psyllas J, Aspe R, Maughan P, Murphy A. Interpreting magnitude of change in strength and conditioning: effect size selection, threshold values and Bayesian updating. J Sports Sci. 2022. doi:10.1080/02640414.2022.2128548 — read at source; pooled thresholds 0.16 [0.15–0.18], 0.46 [0.45–0.48] and 0.81 [0.79–0.83] from 643 studies.
  7. T1 Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863 — read at source; it states the textbook 0.2 / 0.5 / 0.8 thresholds and calls them arbitrary. We cite it rather than Cohen's 1988 book because it is a source you can open and check.
  8. T2 Cochrane Handbook for Systematic Reviews of Interventions, chapter 3: Defining the criteria for including studies. training.cochrane.org (retrieved 19 August 2026) — the source for the split between inactive and active comparisons, for surrogate outcomes, and for setting time points in advance. A methods handbook from an institute, not a study.
  9. T1 Mesquida C, Murphy J, Lakens D, Warne J. Replication concerns in sports and exercise science: a narrative review of selected methodological issues in the field. R Soc Open Sci. 2022;9:220946. doi:10.1098/rsos.220946 — the worked example behind the twenty-versus-forty-four comparison in section 2 comes from this review.

Found a mistake in this piece? Tell us at info@enduranceproof.com — you do not need to be a scientist, and "this number looks off" is a perfectly good message. Anything we correct, and when, goes on our corrections page.

This is educational material about how research is done, not training or medical advice. Nothing here is a recommendation about what any individual should do.

One email when the next piece lands.

A short note when a new check or explainer goes up. No spam, never shared, one click to leave.

No spam, never shared, one click to leave.