Your Durability Test Is Mostly Noise — How to Run One That Means Something
Most published durability tests were never checked for reliability once riders were actually fatigued. Here's the evidence on why, and a protocol built to survive that scrutiny.
Sprint Summary
The short version — read this if you're short on time.
A durability number is only as good as the test that produced it, and most published tests were never checked for reliability in the state that actually matters: fatigued. When researchers have checked, the change scores that durability testing depends on are often barely better than noise at detecting a real 4–5% shift. None of that makes durability testing worthless; it means the rigour is in the design, not the concept. Anchor your preload to an intensity domain and a fixed kJ/kg dose, standardise fuelling and environment, resist the two-point critical-power shortcut, and repeat before you believe a change. If you'd rather add nothing to your training, heart-rate decoupling on a ride you already do gives you a genuinely useful, if imperfect, signal — track it over months rather than trusting any single ride. The simple test in “Fresh Watts vs. Hour-Four Watts” is still a fine place to start; this is what makes the number you get from it worth acting on.
A maximal effort stacked onto two or more hours of prior work carries real heat, dehydration and hypoglycaemia risk. Don't attempt a fatigued maximal test fasted, in the heat, or unaccompanied on the road, and do any 3-minute all-out effort indoors or on a closed climb, never in traffic.
Anyone with cardiovascular risk factors should consult a qualified professional before undertaking repeated maximal fatigue testing.
Almost all of the reliability and protocol evidence here comes from small, male, trained-to-professional cycling samples. Treat specific numbers as directional, not as a personal guarantee, and note the triathlon-specific evidence is limited to one small study.
Full Distance
The complete research and analysis.
Most published durability tests are less scientific than they look. You do a hard preload, then a maximal effort, and you get back a single percentage — “lost 7% after three hours” — and you treat it like a lab result. The honest picture, once you look at how these protocols were actually validated, is that a lot of that number is noise, not signal.
We've already covered a simple at-home version of this test in “Fresh Watts vs. Hour-Four Watts.” That article gave you a way to see whether your watts fade late in a ride. This one is about the layer most riders skip: whether the number you got back is trustworthy enough to act on, and what changes if it isn't.
Why an unreliable number is worse than no number
If you train 6–10 hours a week around a job, you don't get many chances to run a big durability test. If a result moves you to change your training — more long rides, less intensity, a different fuelling strategy — based on a number that's mostly test-to-test variation, you've spent a limited resource, your training time, reacting to noise.
This matters more for durability tests than for most fitness tests, because the whole point of the test is to detect a change between two states, fresh and fatigued, and change scores are consistently the least reliable numbers in this literature.
What the evidence says
The reliability problem nobody flags
A 2025 scoping review of cycling race-simulation and durability protocols screened 3,114 records and settled on 48 included articles describing 30 unique protocols. Its finding: nineteen of those thirty studies did not demonstrate or provide any data on the reliability of the post-preload performance test — the exact test being used to declare someone's durability had improved or declined Peeters, Barrett & Podlogar, European Journal of Applied Physiology, 2025. A further four protocols borrowed reliability figures from earlier studies run under different conditions, two cited unpublished lab data of unclear state, and three measured reliability only in the fresh state — which tells you nothing about how noisy the test is once you're actually tired. Only one protocol in the entire review, a 6 kJ per kilogram body-mass test, had test–retest reliability established after the relevant pre-load, with a coefficient of variation of 2.4%, as reported via the same review, Peeters 2025.
When researchers have gone looking for that reliability directly, the results are not reassuring. In 26 trained runners tested twice on a roughly 2.5-hour run at 89% of speed at the first ventilatory threshold, the absolute fatigued-state values were reasonably reliable (ICC 0.83–0.97) — but the durability scores themselves, calculated as the change between fresh and fatigued, were mostly not: the change in%VO2max had a reliability of ICC 0.36–0.38, the change in running economy ranged from ICC 0.07 to 0.57, and only the change in peak speed reached a reliability most fields would call acceptable (ICC 0.81). The minimal detectable change for these measures — the smallest change you could call real rather than test noise — was roughly 6–12% Malinen, Nuuttila, Matomäki, Uusitalo & Kyröläinen, European Journal of Sport Science, 2026. This is a study of trained runners, not cyclists, but the arithmetic transfers wherever it's applied: if your test moves 4%, and the smallest change that clears the noise floor is 6–12%, you have not learned anything about whether you improved.
The protocol that survives that
The people who built this field have published an unusually explicit set of design rules for making a durability test mean something, and they translate directly into consumer advice Hunter, Maunder, Jones, Gallo & Muniz-Pumares, Experimental Physiology, 2025.
- Anchor the preload to intensity domains — moderate, heavy, severe — not a percentage of VO2max or peak power; the review calls percentage-of-max prescriptions questionable and recommends domains instead.
- Match the work dose in kilojoules per kilogram of body mass within each domain, not time. A fixed-time protocol delivers wildly different absolute work to different riders — in one cited example, 150 minutes at 90% of the first ventilatory threshold delivered roughly double the total kilojoules to the strongest rider compared with the weakest.
- Standardise fuelling ruthlessly, before and during every test. Carbohydrate intake changes the size of the fatigue effect you're trying to measure, so an unstandardised feed can manufacture or erase your entire “durability change.”
- Standardise the environment. Heat and humidity change how much you fatigue independent of the work done, and the review concedes there isn't yet good data on exactly how sensitive durability testing is to environmental conditions.
- Familiarise before you test for real, and standardise the transition from the preload into the test itself — same gap, same setup, every time.
On top of the design rules, the dose itself matters. A verified kilojoule dose–response picture, built from several separate studies, gives an honest sense of what different preloads actually do:
- Under 10 kJ/kg: no significant impairment of power from 10 seconds to 20 minutes, in a field sample of 42 female and 42 male professionals Journal of Science and Medicine in Sport, 2025.
- 2.5 kJ/kg accumulated above critical power: critical power down 2.2%, in 17 male professionals; the same amount of work below critical power produced no change Mateo-March et al., Journal of Science and Medicine in Sport, 2024.
- 6 kJ/kg above critical power: 20-minute power down 13.6%, in 90 male U23 riders Mateo-March et al., Scandinavian Journal of Medicine & Science in Sports, 2026.
- 7.5 kJ/kg above critical power: critical power down 16.2%, in the same 17 male professionals Mateo-March et al., 2024.
- Around 13 kJ/kg (roughly 1,000 kJ at 75 kg): 20-minute power down 6.5% in more successful amateur racers and 12.5% in less successful ones, in 14 amateurs Barsumyan, Soost & Burchard, Frontiers in Sports and Active Living, 2025.
- 40 kJ/kg (about 4 hours): 20-minute time-trial power down 2.9% on average, with an individual range from −8.5% to +1.1%, in 12 male professionals Valenzuela et al., International Journal of Sports Physiology and Performance, 2023.
The clearest single lesson from that list, and one worth repeating because it contradicts how most riders track “kJ burned,” is that how the work is accumulated matters more than how much. Work above critical power degrades critical power and maximal mean power; work-matched work below critical power does not move either Mateo-March 2024. That's one field study in 17 male professionals and needs replication in amateurs, but it's the reason a durability test built on gentle steady riding and one built on hard intervals are not measuring the same thing, even at an identical kJ total.
Once the preload is dosed properly, the test itself matters. The best-validated single option — the only durability performance test with published reliability actually measured in the fatigued state — is the fatigued 3-minute all-out test. After 2 hours of heavy-intensity cycling, its end-power estimate was near-identical on repeat trials (273 ± 52 W vs 276 ± 58 W, r = 0.99), and it agreed closely with a full multi-trial critical-power estimate done the same way (256 ± 41 W vs 256 ± 52 W) Clark, Vanhatalo, Bailey et al., Medicine & Science in Sports & Exercise, 2018; Clark et al., American Journal of Physiology–Regulatory, Integrative and Comparative Physiology, 2019. Both studies used small, male-only samples (n = 6, 9 and 14), so treat the magnitude as moderate-to-strong evidence rather than settled fact — but it's the strongest reliability evidence in the field.
One trap to avoid entirely: estimating critical power from just two efforts. The methodological review is blunt about why — a two-point line will always fit perfectly, so there's no way to judge how good that fit actually is, and if your shorter effort fades more than your longer one under fatigue, the two-point method can push your estimated critical power upward, exactly backwards from what happened Hunter 2025. Several popular apps build their fatigue-resistance numbers on exactly this method.
The test you're already doing: heart-rate decoupling
If you don't want to add a dedicated maximal test at all, there's a passive option with genuinely good — though incomplete — evidence behind it: heart-rate decoupling on a standard long ride.
Pooling four studies covering 51 trained cyclists, heart-rate decoupling correlated strongly with the loss of power at the aerobic threshold (r = −0.76). A model combining baseline threshold power with the change in respiratory frequency, the change in heart rate over time, and peak aerobic capacity predicted real-time threshold power to within about 7 watts on average (R² = 0.95; a bootstrap check gave 8.3 W and R² = 0.93). The authors conclude decoupling “may be a practical method of durability assessment” Rothschild, Gallo, Hamilton et al., European Journal of Applied Physiology, 2025. Graded fairly, this is moderate evidence: a strong association in a pooled sample, with a prediction model that hasn't yet been tested outside that sample.
The caveats matter as much as the headline number. There is no agreed value for what counts as “too much” decoupling in cycling, and the methodological review is explicit that “limited research has demonstrated conflicting findings regarding the relationship between HR decoupling and deterioration of rested physiological parameters” — it remains unclear which decoupling marker to use, and what magnitude should trigger concern Hunter 2025. Drift is also strongly condition-dependent: it averaged just 2.09 ± 2.29% over an hour at 75% FTP indoors with a directed fan, standardised fluids and no food, in 17 trained cyclists Barsumyan, Soost, Graw & Burchard, BMC Sports Science, Medicine and Rehabilitation, 2026 — but heart rate rose 12% over just 45 minutes at 35°C in a separate study of 9 male cyclists Wingo, Lafrenz, Ganio, Edwards & Cureton, Medicine & Science in Sports & Exercise, 2005. Comparing your decoupling number on a hot day with your number on a mild one is comparing two different tests.
There's a useful, free tell buried in the same cadence data: cadence decline is mechanically linked to drift. In 17 trained cyclists tested monthly for five months on a standard 60-minute effort, cadence fell an average of 1.75 rpm, and every 1-rpm drop was associated with an extra 0.61 percentage points of cardiovascular drift Barsumyan et al., 2026. If your cadence is quietly sliding on a ride that's supposed to feel steady, your decoupling number is probably sliding with it — worth a look before you even check heart rate.
The one population-scale anchor for a decoupling threshold comes from running, not cycling: across 82,303 marathon finishers, a decoupling value of 1.025 distinguished runners who held their pace from those who faded Smyth, Maunder, Meyler, Hunter & Muniz-Pumares, Sports Medicine, 2022. That's the closest thing to a validated number in this space, and it's still not a cycling number.
How to actually run this
- If you want a maximal test: use a fixed kJ/kg preload matched to an intensity domain, not a percentage of your max, standardise fuelling and environment, and finish with either a 20-minute effort or, for insight into critical power and W′ specifically, a genuine 3-minute all-out effort. Do it twice, at least two weeks apart, before you believe the number.
- Treat any change under about 5% as noise unless you've repeated the test and seen the same direction — published minimal detectable changes for durability measures run from roughly 6% to 12% Malinen 2026; Mateo-March et al., International Journal of Sports Medicine, 2025.
- If you'd rather add zero extra training stress: pick one existing long ride a month, keep the route, fuelling and, as far as you can, the temperature identical, and track heart-rate decoupling over months rather than trusting any single ride.
- If you're a triathlete and want a cycle-then-run version: the closest thing to a validated protocol used a preload of 20 kJ/kg (men) or 15 kJ/kg (women) at 85% of the second lactate threshold, followed by an incremental run, and found good reliability (ICC ≥ 0.845, typical error ≤ 3.6%) for most outcomes in 10 triathletes — except fractional utilisation at the second threshold, which was unreliable (ICC 0.193) Keller, Röhrs & Wahl, International Journal of Sports Physiology and Performance, 2026. It's a small study (n = 10), but it's the only durability protocol in this research built around the run, not just the bike.
One honest limitation runs through nearly everything above: almost all of it comes from male cyclists tested on the bike alone. If you're a triathlete, a number that describes your watts at hour four says nothing directly about what your legs will do on the run — that's a genuinely different, much thinner evidence base.
Common mistakes
- Testing with a one-minute all-out effort after the preload. This is the most-blogged version of a durability test, but short efforts have the worst field repeatability of any duration measured — 10-second power was the least reliable interval in a large field dataset — making a 1-minute test the noisiest choice available, not the simplest Mateo-March et al., International Journal of Sports Medicine, 2025. Coaching content that recommends a 1-minute post-preload test, including a widely shared TrainingPeaks fatigue-resistance article, is a practitioner opinion, not a validated method.
- Letting an app estimate critical power from two efforts and calling that your fatigued critical power. As above, this method can push the number in the wrong direction under fatigue.
- Changing your fuelling, route, or the weather between a test and its re-test, then reading the difference as durability.
- Using a fixed ride time and assuming that's a fixed work dose. In one controlled 90-minute bout, men accumulated 17.56 ± 2.82 kJ/kg and women 14.08 ± 2.25 kJ/kg — a real difference in the actual stimulus delivered by “the same” session Pastorio et al., Scandinavian Journal of Medicine & Science in Sports, 2026.
- Acting on a single test result. Given the change-score reliability numbers above, one test tells you almost nothing; a repeated test, done the same way, starts to tell you something.
How to apply this week
- Pick the long ride you'd have done anyway. Note your body mass and estimate the kJ/kg you're about to accumulate, so you have a dose to compare against next time.
- Keep fuelling and, as far as you can control it, the route and time of day consistent — write down what you did so you can copy it exactly.
- If you want a maximal component, finish at a standard accumulated kJ/kg with either a 20-minute effort or a 3-minute all-out effort, done somewhere safe — indoors or on a closed climb, never in traffic.
- If you'd rather not add a maximal effort, track heart-rate decoupling for that ride on whatever platform you already use, and glance at your cadence at the same time.
- Put a note in your calendar to repeat the identical session in 2–4 weeks. One number is a data point; two numbers, done the same way, start to be a test.
References
- Peeters, Barrett & Podlogar, “What is a cycling race simulation anyway: a review on protocols to assess durability in cycling,” European Journal of Applied Physiology, 2025
- Malinen, Nuuttila, Matomäki, Uusitalo & Kyröläinen, “Test-Retest Reliability of Physiological Resilience During and After Prolonged Moderate-Intensity Running in Well-Trained Runners,” European Journal of Sport Science, 2026
- Hunter, Maunder, Jones, Gallo & Muniz-Pumares, “Durability as an index of endurance exercise performance: Methodological considerations,” Experimental Physiology, 2025
- Mateo-March, Leo, Muriel, Javaloyes, Mujika, Barranco-Gil, Pallárés, Lucia & Valenzuela, Journal of Science and Medicine in Sport, 2024
- Mateo-March et al., Scandinavian Journal of Medicine & Science in Sports, 2026
- Sex differences in durability: A field-based study in professional cyclists, Journal of Science and Medicine in Sport, 2025
- Barsumyan, Soost & Burchard, “Enhanced durability predicts success in amateur road cycling,” Frontiers in Sports and Active Living, 2025
- Valenzuela, Alejo, Ozcoidi, Lucia, Santalla & Barranco-Gil, “Durability in Professional Cyclists: A Field Study,” International Journal of Sports Physiology and Performance, 2023
- Clark, Vanhatalo, Bailey, Wylie, Kirby, Wilkins & Jones, Medicine & Science in Sports & Exercise, 2018
- Clark et al., American Journal of Physiology–Regulatory, Integrative and Comparative Physiology, 2019
- Rothschild, Gallo, Hamilton, Stevenson, Dudley-Rode, Charoensap, Plews, Kilding & Maunder, European Journal of Applied Physiology, 2025
- Barsumyan, Soost, Graw & Burchard, BMC Sports Science, Medicine and Rehabilitation, 2026
- Wingo, Lafrenz, Ganio, Edwards & Cureton, Medicine & Science in Sports & Exercise, 2005
- Smyth, Maunder, Meyler, Hunter & Muniz-Pumares, Sports Medicine, 2022
- Keller, Röhrs & Wahl, International Journal of Sports Physiology and Performance, 2026
- Mateo-March et al., International Journal of Sports Medicine, 2025
- TrainingPeaks, “Race Stronger, Longer: How to Build Fatigue Resistance,” 2025 (practitioner opinion)
- Pastorio et al., Scandinavian Journal of Medicine & Science in Sports, 2026
Frequently asked questions
How much does my durability test result need to change before I should trust it?
Published minimal detectable changes for durability change-scores are roughly 6–12%, so treat anything smaller as noise unless you've repeated the test and seen the same direction (Malinen et al., 2026).
What's the single most reliable durability test available?
The fatigued 3-minute all-out test is the only durability performance test with published reliability actually measured after a real fatigue preload, agreeing on repeat trials with r = 0.99 (Clark et al., 2018/2019).
Should I trust the fatigue-resistance number my training app calculates?
It depends how it's built. If it estimates critical power from just two efforts, be cautious — that method can bias the fatigued critical-power estimate upward instead of down (Hunter et al., 2025).
Is heart-rate decoupling a good enough substitute for a maximal test?
It's a reasonably good signal (r = −0.76 against threshold-power loss in pooled data), but there's no agreed decoupling threshold and drift is highly condition-dependent, so compare rides done in similar conditions only (Rothschild et al., 2025; Hunter et al., 2025).
Does a fixed ride time guarantee a fixed durability dose?
No. In one 90-minute bout, men accumulated more work per kilogram than women at the same duration, so dose your preload in kJ/kg, not minutes (Pastorio et al., 2026).
Is there a durability test built for triathletes specifically?
One small study (n = 10) validated a cycle-then-run version with good reliability for most outcomes, though one specific measure was unreliable — it's the closest thing available, but still a small sample (Keller, Röhrs & Wahl, 2026).
Get The Forward
Research-led triathlon guidance, straight to your inbox. No spam, no filler — unsubscribe any time.
