A laboratory number is real, and it is usually not the number that will show on the clock. Part of that gap is arithmetic: above roughly jogging pace, every per cent of energy you save buys you less than a per cent of speed. Part of it is deeper — a measurement of capacity and a race result are different quantities, and one can move while the other stays where it was. Everything the road throws at you sits on top of those two, not underneath them.
- One number, two questions, two answers. Across a group of people, a laboratory measure of aerobic capacity lined up closely with how they performed. Within those same people, after six weeks of training, how much that measure improved said nothing about how much their performance improved.
- The exchange rate is not one to one, and it moves. At world-class marathon pace about two thirds of an energy saving shows up as speed. At jogging pace slightly more than all of it does.
- Knowing the rate makes the number useful. The four per cent measured for a carbon-plated racing shoe, converted properly, predicted a marathon of 2:01:36. The world record at the time was 2:01:39.
- The laboratory number wobbles too. Between your first and second attempt at a test you improve by about 1.2 per cent on average, and afterwards by 0.2 per cent. The first of those is learning the test.
A study comes out. Some intervention — a shoe, a supplement, a block of intervals — improved something by four per cent in a laboratory. You do the arithmetic on your own marathon: four per cent of three and a half hours is about eight minutes. That is a lot of minutes.
It is also almost certainly too large, and the interesting part is why. The obvious explanation is that a race is messy and a laboratory is not: there is wind out there, and hills, and you have to pace yourself, and you are tired by thirty kilometres. All of that is true, and it is the smallest of the three reasons.
1 The same number answered two opposite questions
In 2009 a group of researchers put twenty-four untrained men through six weeks of supervised cycling and measured them carefully before and after.1 Two of their measurements matter here: maximal oxygen uptake, the classic laboratory measure of aerobic capacity, and a fifteen-minute time trial on a bicycle — as hard as you can go for a quarter of an hour.
Before the training started, those two lined up well. Knowing someone's oxygen uptake told you a great deal about how they would do in the time trial. In the language of the paper, the relationship accounted for about eighty per cent of the differences between people.
Then everyone trained for six weeks, and the researchers asked the second question: did the people whose oxygen uptake improved most also improve most in the time trial? The answer in their words is that the change in maximal oxygen uptake “was completely unrelated to the change in aerobic performance”. Figure 2 sets one against the other.
Read those two findings next to each other and the usual explanation stops working. The time trial was done in the laboratory. There was no wind in it, no hills, no tactics, nobody going off too fast in the first kilometre. Both measurements were taken on the same equipment, in the same room, on the same people. If the distance between a laboratory number and a performance were mainly about the mess of the outside world, this gap could not have opened up where it did.
2 A capacity and a result are different quantities
Maximal oxygen uptake is a ceiling: the most oxygen your body can take in and use, measured while you are driven to exhaustion under one fixed instruction. A race is not a ceiling (figure 3). It is how fast you covered a distance, which depends on that ceiling and on how efficiently you move and on how much of the ceiling you can hold for an hour and on how well you judged your effort.
That is simply what the measurement is. A review of the physiology of elite endurance athletes builds the whole picture out of three laboratory quantities — the ceiling, the point at which your effort stops being comfortably sustainable, and how much energy your movement costs2 — and those three genuinely do describe who is fast. An expert panel asked to list what determines endurance performance came back with twenty-six factors.3 Those three are on the list. So are twenty-three other things.
So when a study reports that an intervention improved a laboratory measure, it has told you something true about a capacity. Whether that capacity was the thing holding you back is a separate question, and one the study usually has not asked.
The field has known this for a long time and has said it plainly. A 1999 paper on how to design research into performance enhancement sets out the conditions a test has to meet before a change in it can stand in for a change in an event — it has to resemble the event, wobble less than the event does, and let you estimate the improvement in the event. Where those conditions are not met, the authors write, “the event itself provides the only dependable estimate of performance enhancement”.13 That is a strong sentence, and it comes from inside the discipline rather than from anybody complaining about it.
This is not a problem confined to oxygen uptake. A systematic review of branched-chain amino acid supplements set out specifically to ask whether the biochemical changes these supplements produce turn into performance, found the biochemical changes, and found no consistent improvement in performance.4 When a national Olympic committee's test battery for alpine skiers was checked against actual competition results, the physiological tests could not predict them at all — the authors' conclusion was that tests should be validated against competition before anybody plans training around them.5 That last one is a different sport and it tells you nothing quantitative about endurance running. It tells you the pattern is not a quirk of one paper.
3 So how much of a gain does arrive?
Sometimes all of it and more, sometimes two thirds, and which one depends on how fast you are already going. This part is pure arithmetic, and it is worth following because it is the piece that most people have never been shown.
Suppose something genuinely lowers the energy it costs you to run at a given speed — better shoes, a better stride, a lighter kit. You now have energy to spare. How much faster can you go on it?
The tempting answer is: four per cent cheaper, so four per cent faster. That would be right if energy cost rose in a straight line with speed. It does not. Oxygen uptake climbs more steeply the faster you run, and on top of that, pushing air out of the way costs energy in proportion to the cube of your speed6 — so at racing speeds a large share of any saving is eaten by simply going faster.
Two things in that curve are worth keeping. The first is that for most people reading this, the exchange rate is favourable: at a comfortable jogging pace, a ten per cent energy saving is worth about twelve and a half per cent more speed, because at that pace air resistance is barely charging you anything. The second is that the faster you already are, the less of any saving you keep. At the pace a world-class marathon is run, about two thirds of it arrives.
4 The shoes, from laboratory to finish line
In 2017 a group of researchers measured eighteen strong runners in a prototype racing shoe with a compliant foam midsole and a stiff embedded plate. Averaged across three speeds, it cost about four per cent less energy to run in than two established marathon shoes.7 They ended the paper by predicting that top athletes in these shoes could run substantially faster.
Put that four per cent through the exchange rate above and you can be specific about “substantially”. Starting from a 2:04 marathon, a three per cent improvement in running economy works out at 1.97 per cent more speed, which is a finishing time of 2:01:36.6
The world record at the time that calculation was published stood at 2:01:39. Three seconds over forty-two kilometres.
That is the part worth holding on to. The laboratory number was sound and the extrapolation was careful. The measurement was good, the conversion was known, and the answer was close enough to be slightly uncanny. What would have been wrong is the arithmetic most people do in their heads. Take the same three per cent straight off the finishing time and you get 2:00:16 — one minute and twenty-three seconds faster than the record that actually stood, and faster than anybody has run a record-legal marathon to this day.
5 And then there is everything a laboratory does not have
This is the layer everybody thinks of first, and it is real. It is simply the third one rather than the first.
A treadmill in a room has still air, a surface that does not change, a speed chosen for you, and a runner who arrived rested. A race has wind, corners, camber, hills, weather, a pace you have to judge yourself, and a body that has already been working for two hours by the time the interesting part starts. None of that is in the measurement, and figure 6 lists the ones that matter most.
What makes this the third layer rather than the first is the order of size. The arithmetic in section 3 takes a third off an energy saving at racing pace before the weather is mentioned, and the distinction in sections 1 and 2 can take all of it. Race-day conditions then move the result again, in a direction that depends on the day.
6 Even the bridge between treadmill and road wobbles
Laboratories know their treadmills are not the road, and they have a standard correction for it. In 1996 two researchers ran nine trained runners at six speeds, on the treadmill at several gradients and once outdoors on a level road, and found that at a one per cent gradient the energy cost on the belt matched the energy cost outside.8 A one per cent incline compensates for the air you are not pushing through.
That finding has been standard equipment in the field ever since. In 2026 a review looked back over the thirty years of evidence that accumulated after it and concluded that the energetic equivalence between treadmill and overground running is influenced by running speed, environmental conditions, methodological factors and differences between individuals — so a fixed one per cent may not consistently reflect the cost of running outside across different contexts or populations.9
Take that the right way round. A 1996 paper offered a practical correction, the field adopted it, and three decades of further measurement has now mapped where it holds and where it does not. That is not a failure; figure 7 is what it looks like written out. It does mean that the bridge a laboratory uses to talk about the road is itself an approximation with conditions attached, and those conditions are usually not stated in the sentence you read in a headline.
7 The laboratory number does not sit still either
Measure the same athlete twice and you get two numbers. How far apart they are depends on which test you chose.
A time trial — go as hard as you can over a set distance — typically varies by under five per cent from attempt to attempt. A test where you hold a fixed intensity until you cannot continue varies by more than ten.10 That matters for what a study can detect: the difference between winning and coming second in a real event is often under one per cent, so a test that wobbles by ten cannot see an effect worth having. Figure 8 shows that contrast. On a bicycle, tests where the rider chooses their own pace carry a random error of about two to three per cent, and the best protocols get down to about one.11 None of this is obscure: thirty-three specialists asked to agree on what a performance test should be judged on produced a checklist, and how repeatable the test is sits on it.14
There is a quieter version of the same problem. A meta-analysis of performance-test reliability found that between a first and second attempt at a test, people improve by about 1.2 per cent, and between later attempts by about 0.2 per cent.12 The first of those is learning the test. A study that does not give its participants a practice run hands that 1.2 per cent to whatever it was testing.
How large a difference has to be before it is worth anything to you is its own question, and it has a real answer that depends on your level — a half of one per cent decides an Olympic final and is invisible to someone running a four-and-a-half-hour marathon. That is a piece we have not written yet.
8 So what does this leave you with?
Something other than scepticism about laboratories. Something more useful than that: the three questions in figure 9, which turn a headline percentage back into a quantity you can reason about.
- Better at what? A study that measured a finishing time and a study that measured a physiological quantity have done different things. Both are worth reading. Only one of them has told you about a result.
- Measured where, and at what pace? On a belt in still air, a saving is worth more than it is on the road, and the gap widens the faster the pace. A percentage quoted without a speed attached is missing the information required to convert it.
- In whom? Effects found in untrained people frequently shrink in trained ones, and the conversion from energy to speed differs by pace, which means it differs by athlete. The exchange rate that applies to a two-hour marathon runner is not the one that applies to you.
A laboratory number is real currency. It is just not the same currency as the clock, and nobody prints the exchange rate on the headline.
How we did this (for the curious)
This is an explainer rather than a review, so there is no systematic search behind it. Before writing we wrote a scoping document: what the subject is, what belongs in it, what does not, and where the sources disagree. We built that from fourteen sources chosen to cover measurement standards, definitions and consensus rather than to estimate an effect, and it is what this piece was then held to.
Five of those fourteen we hold in full. For the other nine we read the abstract at source and have taken from them only what the abstract states — no derivations, no extrapolations. That distinction matters most for the 2009 cycling study in section 1, which carries the opening finding and which we could not obtain in full: the publisher's embargo covers the repository copy as well, and five routes failed. Both of the sentences we use from it appear verbatim in its abstract, but we have not read its discussion, and section 1 says so.
One number on this page is our own calculation rather than a source's: the curve in figure 4, and the conversion in figure 5 that follows from it. We rebuilt the published equation in code instead of reading values off the paper's chart, and checked the result against the numbers the authors state. It reproduces their worked example exactly and sits about two to three tenths of a percentage point from two other anchors they give, which we could not account for. That is why the text says “about two thirds” rather than a precise figure.
What we left out deliberately: how large a difference has to be before it is worth having, which belongs with a separate piece; what a threshold is and how to find yours, which has its own piece coming; what heat and humidity do to your pace; and re-examining the carbon-shoe claim itself, which already has its own review. The shoes appear here as a worked example, not as a subject.
Sources
Fourteen sources. They were chosen to establish what the subject is and how it is measured, not to estimate the size of any effect. All are research (T1). Five we hold in full and read in full; for the nine marked abstract we read the abstract at its own address and use only what it states. After each entry: what we took from it.
- T1 Vollaard NBJ, Constantin-Teodosiu D, Fredriksson K, Rooyackers O, Jansson E, Greenhaff PL, Timmons JA, Sundberg CJ. Systematic analysis of adaptations in aerobic capacity and submaximal energy metabolism provides a unique insight into determinants of human aerobic performance. Journal of Applied Physiology. 2009;106(5):1479–1486. doi:10.1152/japplphysiol.91453.2008 — abstract — twenty-four sedentary men, six weeks of supervised cycling, a fifteen-minute time trial; maximal oxygen uptake and performance related at baseline (r² = 0.80), and the change in maximal oxygen uptake “completely unrelated to the change in aerobic performance”.
- T1 Joyner MJ, Coyle EF. Endurance exercise performance: the physiology of champions. The Journal of Physiology. 2008;586(1):35–44. doi:10.1113/jphysiol.2007.143834 — abstract — the three-factor account of endurance performance (maximal oxygen uptake, the lactate threshold and efficiency), how they interact, and the authors' own caution that motivational and social factors sit outside it.
- T1 Konopka MJ, Zeegers MP, Solberg PA, Delhaije L, Meeusen R, Ruigrok G, Rietjens G, Sperlich B. Factors associated with high-level endurance performance: An expert consensus derived via the Delphi technique. PLOS ONE. 2022;17(12):e0279492. doi:10.1371/journal.pone.0279492 — eighteen international experts over three rounds arriving at twenty-six factors for endurance performance, of which the familiar laboratory measures are three.
- T1 Del Guerra GC, Ohannesian VA, Semerdjian R, et al. Branched-chain amino acid supplementation and endurance performance: reporting guidelines and systematic review of biochemical vs clinical evidence. The Physician and Sportsmedicine. 2026;54(2):180–191. doi:10.1080/00913847.2026.2627863 — abstract — a review asking specifically whether biochemical changes translate into functional benefit: fifteen studies, consistent biochemical change, no consistent improvement in performance, fatigue or recovery.
- T1 Nilsson R, Theos A, Lindberg AS, Ferguson RA, Malm C. Lack of Predictive Power in Commonly Used Tests for Performance in Alpine Skiing. Sports Medicine International Open. 2021;5(1):E28–E36. doi:10.1055/a-1078-1441 — fourteen elite female alpine skiers; the national test battery produced no valid model for competitive ranking, and the authors' conclusion that test batteries should be validated against competition before training is planned around them. A different sport, used here only for the pattern.
- T1 Kipp S, Kram R, Hoogkamer W. Extrapolating Metabolic Savings in Running: Implications for Performance Predictions. Frontiers in Physiology. 2019;10:79. doi:10.3389/fphys.2019.00079 — the oxygen uptake–velocity equation used for figure 4 and the conversion in figure 5; the cubic cost of air resistance; that below about 3 m/s a saving returns more than itself and above it less; about two thirds at 5.5 m/s; and the worked example in which a three per cent economy gain from 2:04 pace gives 1.97 per cent and 2:01:36.
- T1 Hoogkamer W, Kipp S, Frank JH, Farina EM, Luo G, Kram R. A Comparison of the Energetic Cost of Running in Marathon Racing Shoes. Sports Medicine. 2018;48(4):1009–1019. doi:10.1007/s40279-017-0811-2 — eighteen high-calibre runners at 14, 16 and 18 km/h; the prototype shoe cost 4.16 and 4.01 per cent less energy than two established marathon racing shoes with mass matched, in all eighteen runners, independent of speed.
- T1 Jones AM, Doust JH. A 1% treadmill grade most accurately reflects the energetic cost of outdoor running. Journal of Sports Sciences. 1996;14(4):321–327. doi:10.1080/02640419608727717 — abstract — nine trained male runners at six velocities between 2.92 and 5.0 m/s, on the treadmill at five gradients and once outdoors on a level road; the one per cent gradient matching road running across the middle of that range.
- T1 Bottura R, Fletcher J. Beyond the 1% treadmill rule: 30 years of evidence on treadmill and overground running economy. Applied Physiology, Nutrition, and Metabolism. 2026;51:1–7. doi:10.1139/apnm-2026-0109 — abstract — a re-examination of the fixed one per cent correction, concluding that equivalence is influenced by speed, environment, methodological factors and individual variability, and advocating individualised and context-specific approaches when extrapolating from laboratory to outdoors.
- T1 Currell K, Jeukendrup AE. Validity, reliability and sensitivity of measures of sporting performance. Sports Medicine. 2008;38(4):297–316. doi:10.2165/00007256-200838040-00003 — abstract — the three properties of a performance test; time-to-exhaustion protocols varying by more than ten per cent against under five for time trials; and that the difference between first and second place in a sporting event is often under one per cent.
- T1 Paton CD, Hopkins WG. Tests of cycling performance. Sports Medicine. 2001;31(7):489–496. doi:10.2165/00007256-200131070-00004 — abstract — random error of at least about two to three per cent in tests requiring the rider to choose their own pace, and as low as about one per cent for the most controlled protocols.
- T1 Hopkins WG, Schabort EJ, Hawley JA. Reliability of power in physical performance tests. Sports Medicine. 2001;31(3):211–234. doi:10.2165/00007256-200131030-00005 — abstract — a meta-analysis of 101 studies; variation between the first two trials 1.3 times that between later ones, with performance improving by 1.2 per cent between the first two and 0.2 per cent between subsequent ones; and larger variation in non-athletes than athletes.
- T1 Hopkins WG, Hawley JA, Burke LM. Design and analysis of research on sport performance enhancement. Medicine & Science in Sports & Exercise. 1999;31(3):472–485. doi:10.1097/00005768-199903000-00018 — abstract — that enhancements in a test and in an event may differ when the factors affecting performance differ between them, and that where a test does not meet its conditions “the event itself provides the only dependable estimate of performance enhancement”.
- T1 Robertson S, Kremer P, Aisbett B, Tran J, Cerin E. Consensus on measurement properties and feasibility of performance tests for the exercise and sport sciences: a Delphi study. Sports Medicine – Open. 2017;3:2. doi:10.1186/s40798-016-0071-y — thirty-three subject-matter experts over two rounds; the checklist of measurement and feasibility properties a performance test should be assessed on, which is the nearest thing this area has to a standard.
Found a mistake here? Tell us at info@enduranceproof.com — you do not need to be a scientist, and "this number looks off" is a perfectly good message. Anything we correct, and when, goes on our corrections page.
This is educational material about how performance is measured and how laboratory results relate to competition, not training or medical advice, and not a recommendation for or against any method, product or protocol.