You cannot check from your kitchen table whether a study is true. You can work out how far it travels — and that is nearly always the more useful thing to know. Seven questions get you there, and all seven are answered in the paper itself.
- Six of the seven sit in the methods section — the part most readers skip on the way to the conclusion. It is the most informative page in the paper.
- Training studies are small — half of them in one flagship journal had nineteen people or fewer. That is normal, and it changes what you can fairly ask of a single paper.
- "Compared with what" is the question that goes missing most often, because a headline has room for only half the sentence.
- Around eight in ten published sport studies report a finding — which tells you something about how papers get published, and is worth knowing before you count them up.
Prefer to watch? All seven questions in 5:09, narrated, with subtitles you can switch off — it has its own page too. Every number in it is traced to a source below.
You read that runners improved three percent after eight weeks of something. The first question that comes to mind is usually is that true? — and that one is close to unanswerable from where you are sitting. There is a better question, and the paper will answer it for you: how was this measured?
A training study is assembled out of a small number of decisions. Who took part. How many of them. How long it ran. What the other group did instead. What got measured, and how big the difference turned out to be. Those decisions are what a finding is made of, and they decide how far it travels — whether it is about people like you, over a time span like yours, against the alternative you were actually considering.
They are also written down, nearly always in the same order, and none of them require any statistics to understand. What follows is where each one lives and what to make of it.
1 Who was in it?
The methods section opens by telling you who this finding is about, and that single paragraph decides most of what the study can mean for you.
The detail that moves the needle furthest is training status. Somebody who has never trained has a great deal of room to improve, and in that state almost any sensible programme produces a result. The closer someone is to what their body can do, the smaller the remaining gains and the harder they are to find. A study run on untrained volunteers is a perfectly good study; it is just answering a question about untrained volunteers.
The rest of the paragraph is worth the same thirty seconds. Age, sex, and the level people were competing at all shape who the answer belongs to. None of this makes a study weaker — it makes it specific, which is what a study is supposed to be. The reading error is ours, when we quietly assume the participants were people like us.
A concrete one from our own shelf: the famous four-percent figure for carbon-plated racing shoes was measured on eighteen very fast men running at around two-and-a-half-hour marathon pace. That is a real result about those runners at that speed. Whether it is a result about a four-hour marathoner is a different question, and the paper never claimed to answer it. We took that apart in the super-shoes check.
2 How many took part?
Training studies are small. Not carelessly small — lab time is expensive, good athletes are scarce, and asking twenty people to change how they train for two months is already a serious undertaking.
One survey of the Journal of Sports Sciences put the median at nineteen participants (figure 4). A look at four other sport-science journals found averages between fifteen and thirty-two, with a spread wider than the average itself — which is another way of saying that the small studies are not outliers, they are the shape of the field.
What matters is what that buys you. A small study is built to show a large effect. It is poorly equipped to rule out a small one. Give twenty athletes a medium-sized difference to find and the study lands on it fewer than half the times you run it; make the group forty-four and it lands four times in five (figure 5).
That has one very practical consequence, and it is probably the most useful single sentence in this piece. When a small study reports no difference between the groups, that usually means there were too few people to see one — not that there is nothing there. A null result from twenty participants is a shrug, not an answer.
There is a design that makes small studies punch well above their size, and it is worth recognising: the crossover. Instead of comparing one group of people with another, every participant does both conditions and is compared with themselves. That removes the biggest source of noise in sport — the fact that people differ enormously from one another — and it is why a well-run crossover on fifteen runners can be more informative than two groups of thirty.
3 How long did it last?
Where a study stops decides what it is able to tell you, and the answer is usually one line in the methods: eight weeks, twelve weeks, one afternoon.
Two things ride on that number. The first is what could even have changed in the time available. Some adaptations move within a fortnight; others need months of consistent work before they show up at all. A six-week study that reports no change in something that takes half a year to shift has not found an absence — it stopped early.
The second is what happens after the last measurement, which is: nobody knows. Early gains in almost any new programme are partly the novelty of a new stimulus, and part of that fades. Whether what remains keeps building or settles back is a question the study did not ask.
Methodologists take this seriously enough to have a rule about it: decide in advance which time points you care about, so that nobody chooses afterwards which measurement to report. When a paper measured at four, eight and twelve weeks and the headline quotes one of them, it is fair to wonder what the other two showed.
4 Compared with what?
Every result is a difference between two things, and the thing on the other side decides what the difference means. This is the question that disappears most reliably, because a headline has room for only half the sentence.
Methodologists split comparisons into two kinds. An inactive comparison is a placebo, a waiting list, no treatment at all, or simply carrying on as before. An active comparison is a real alternative — a different version of the same thing, or the cheaper option you would otherwise have chosen. Both are legitimate, and they answer different questions. Beating nothing tells you the thing does something. Beating the sensible alternative tells you whether to switch, and that is nearly always the question you actually have.
Sport adds a complication that medicine mostly avoids: you can rarely hide which group someone is in. You know whether you are wearing the stiff plated shoe or the flat one, and whether you got the drink or the water. Where the comparison is visible, part of the difference is expectation — and expectation moves performance by an amount you can measure. That is not a reason to set the study aside. It is a reason to read the comparison as carefully as the headline number.
The carbon-shoe research is a clean example again: the comparison there is usually a light racing flat, itself a fast shoe, rather than an ordinary trainer. The famous saving is a difference between two good shoes. Read against the wrong comparison, it sounds roughly twice as useful as it is.
Three things to look for. What did the other group actually do — nothing, or something plausible? Could either side tell which group they were in? And if the comparison is "usual training", whose usual training is that: yours, or that of the students who volunteered?
5 What was measured?
The outcome a study measures is rarely the outcome you have in mind, and the distance between the two is where a lot of headlines are made.
Take running economy — how much oxygen you burn to hold a given pace. It is a sensible thing to measure: it can be done in an afternoon, it is precise, and it is genuinely related to performance. But it is a stand-in. What you actually want to know is your time on race day, and that takes a season, a start line and a good deal of luck to observe.
Cochrane's methods handbook is blunt about stand-ins of this kind: laboratory results "may not predict clinically important outcomes accurately", and plenty of interventions improve the stand-in without improving the thing it stands in for. The same logic applies to a lactate value, a peak power number or a jump height. None of them are wrong. They are just one or two steps back from the finish line.
One more thing hides in the same paragraph: which outcome the researchers said they would look at. A study that measures fifteen things will almost certainly find something interesting in one of them, and a finding that was chosen after the fact is a much weaker claim than one that was named in advance. Good papers say which was the primary outcome. It is worth checking whether the number in the headline is that one.
6 How big was the difference?
Here is the one place where a word does more damage than any other: significant. In ordinary English it means important. In a paper it means something much narrower — that a difference this size is unlikely to have come from chance alone.
Those two sentences say different things, and the gap runs both ways. A large study can find a difference so small that nobody would notice it, and call it significant with a straight face. A small study can miss a difference that would genuinely matter, and report no significant effect. Neither sentence tells you how big anything was.
For that, look for the effect size — a number that expresses how large the difference is relative to how much people vary anyway. Most papers report one, usually as a value with a letter in front of it, and the thresholds everybody quotes for small, medium and large came out of behavioural science decades ago. Sport now has its own: when one group pooled six and a half thousand effects from strength and conditioning trials, the same three words landed on noticeably smaller numbers (figure 10).
The best question of all, though, is more direct than any of this: how much smaller is the difference than the thing you would notice? Researchers call that threshold the smallest meaningful change — the point below which a difference, however real and however significant, is not worth reorganising your week for. Comparatively few sport papers name one, and noticing that it is missing is itself a useful reading skill.
The version of this you can always do yourself is to translate the percentage into time. A one-percent improvement in a forty-five-minute 10K is about twenty-seven seconds; in a four-hour marathon it is a little over two minutes. Whether that is worth anything is a question about your season, not about the study — but you cannot answer it while the result is still a percentage. (Our pace calculator will do the arithmetic.)
7 Would it have been published anyway?
This is the only one of the seven that the paper in front of you cannot answer, and it is the reason to hold every single study a little more loosely than it holds itself.
Research does not work out very often. Most careful ideas turn out to be wrong, or too small to see, or true only under conditions nobody can reproduce. So you would expect a fair share of published papers to report that they found nothing much. In sports and exercise medicine, a survey of 129 papers found that 82.2% reported a positive result; a wider check across kinesiology journals a year later landed on 81.4%, near enough the same number that the authors call the two indistinguishable (figure 11).
That number is telling you about publishing rather than about training, and it is worth being precise about the mechanism, because it is not bad faith. It works from both ends. Journals prefer a paper with a result in it — it gets read, it gets cited, and a page reporting that nothing happened is a hard sell to an editor. And authors prefer to write up the study that found something, for the same reasons plus one more: the paper is easier to write when there is a story in it. Neither party is doing anything they would be ashamed of. The effect is still that the quiet results end up in a drawer, and the ones you read are the survivors.
This is called publication bias, and it is the reason a single striking paper deserves less weight than it feels like it deserves — you are seeing the studies that made it through the filter, not the ones that were run.
The genuinely good news is that the fix exists and you can watch it working. In a registered report, the question, the method and the analysis are reviewed before any data exist, and the journal commits to publishing whatever comes out. Nobody can file the dull result away, and nobody can reshape the hypothesis to fit the numbers. When one team compared 71 registered reports with 152 ordinary ones in the same journals, the share reporting a positive finding went from 96% to 44% (figure 12).
A handful of sport journals now run the format, and when a paper you are reading is a registered report it says so at the top. That is a real mark of quality, and it costs you nothing to look for it.
8 The seven questions
None of these need statistics, and all of them except the last are answered somewhere in the paper. Together they tell you how far a finding travels, which is almost always more useful than knowing whether it cleared a threshold.
- Who was in it? Trained or untrained, and how much like you. This one moves the answer more than any other.
- How many took part? Small is normal. It means the study is built to show a large effect, and that "no difference" probably means "too few people to see one".
- How long did it last? Long enough for the thing being studied to change? And what happened after the last measurement — which nobody knows.
- Compared with what? Nothing, a dummy, usual practice, or the alternative you would actually have chosen. Only the last one answers your question.
- What was measured? And how many steps sit between that and the outcome you care about. Was it the outcome they said they would measure?
- How big was the difference? In seconds per kilometre, not in significance. Big enough that you would notice?
- Would it have been published anyway if the answer had been "no difference"? A registered report is the one design where you know the answer is yes.
A study is a photograph, not a promise. Read for who was in the picture, how many, how long the shutter was open — and what was standing just outside the frame.
Sources
This is an explainer rather than a review: the sources below were picked to explain how a study is put together, not gathered through a systematic search. Eight of the nine were read and checked at their own source rather than through somebody else's summary. That matters more than it sounds: doing it turned up three numbers on the first version of this page that needed correcting. Where a figure still reaches you second-hand, the entry says so.
- T1 Büttner F, Toomey E, McClean S, Roe M, Delahunt E. Are questionable research practices facilitating new discoveries in sport and exercise medicine? The proportion of supported hypotheses is implausibly high. Br J Sports Med. 2020;54:1365–1371. doi:10.1136/bjsports-2019-101863 — read at source; the 82.2% is checked against it.
- T1 Twomey R, Yingling V, Warne J, et al. The nature of our literature: a registered report on the positive result rate and reporting practices in kinesiology. Commun Kinesiol. 2021;1(3). doi:10.51224/cik.v1i3.43 — read at source; they report 81.43% [75.78–86.3].
- T1 Knudson DV. Authorship and sampling practice in selected biomechanics and sports science journals. Percept Mot Skills. 2011;112:838–844. doi:10.2466/17.PMS.112.3.838-844 — read at source; the four journal averages come from its table 1.
- T1 Abt G, Boreham C, Davison G, Jackson R, Nevill A, Wallace E, Williams M. Power, precision, and sample size estimation in sport and exercise science research. J Sports Sci. 2020;38:1933–1935. doi:10.1080/02640414.2020.1776002 — read at source; the median sample size of 19 is checked against it.
- T1 Scheel AM, Schijen MRMJ, Lakens D. An excess of positive results: comparing the standard psychology literature with registered reports. Adv Methods Pract Psychol Sci. 2021;4(2). doi:10.1177/25152459211007467 — read at source; 96% versus 44% checked against it.
- T1 Swinton PA, Burgess K, Hall A, Greig L, Psyllas J, Aspe R, Maughan P, Murphy A. Interpreting magnitude of change in strength and conditioning: effect size selection, threshold values and Bayesian updating. J Sports Sci. 2022. doi:10.1080/02640414.2022.2128548 — read at source; pooled thresholds 0.16 [0.15–0.18], 0.46 [0.45–0.48] and 0.81 [0.79–0.83] from 643 studies.
- T1 Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863. doi:10.3389/fpsyg.2013.00863 — read at source; it states the textbook 0.2 / 0.5 / 0.8 thresholds and calls them arbitrary. We cite it rather than Cohen's 1988 book because it is a source you can open and check.
- T2 Cochrane Handbook for Systematic Reviews of Interventions, chapter 3: Defining the criteria for including studies. training.cochrane.org (retrieved 19 August 2026) — the source for the split between inactive and active comparisons, for surrogate outcomes, and for setting time points in advance. A methods handbook from an institute, not a study.
- T1 Mesquida C, Murphy J, Lakens D, Warne J. Replication concerns in sports and exercise science: a narrative review of selected methodological issues in the field. R Soc Open Sci. 2022;9:220946. doi:10.1098/rsos.220946 — the worked example behind the twenty-versus-forty-four comparison in section 2 comes from this review.
Found a mistake in this piece? Tell us at info@enduranceproof.com — you do not need to be a scientist, and "this number looks off" is a perfectly good message. Anything we correct, and when, goes on our corrections page.
This is educational material about how research is done, not training or medical advice. Nothing here is a recommendation about what any individual should do.