Fresh Watts vs. Hour-Four Watts: Testing and Training Durability
Durability — how well your power holds up late into a long ride — is emerging as a fourth pillar of endurance fitness. Here's what the evidence says, and how to test it at home.
Sprint Summary
The short version — read this if you're short on time.
Durability — how well your power holds up late into a long ride, rather than how high it is when you're fresh — has real, if still-developing, scientific backing as a genuine fourth pillar of endurance performance, and it maps onto a failure mode almost every long-course triathlete recognises: solid testing numbers, then a fade on the bike and a wrecked run. The framework is credible; the amateur-specific evidence for age-groupers is real but comes from a small sample of 14 athletes, so it should inform your training, not dictate a total overhaul of it. Practically, you don't need a lab to get useful information: a simple fresh-fatigue-fresh power test, or even just watching your heart-rate decoupling trend on long rides, gives you a low-cost, personal durability signal worth tracking over a season — tested carefully, fuelled properly, and never in the final weeks before a race.
Full Distance
The complete research and analysis.
Why Your Fresh Numbers Don't Tell the Whole Story
Most age-group triathletes size up their bike fitness with a handful of fresh, well-rested numbers: 5-minute power, 20-minute power, maybe an FTP from a ramp test. Those numbers are genuinely useful, and they predict a lot about a 40-minute time trial done on fresh legs. What they predict far less well is what happens after two, three or five hours in the saddle on the way into a 70.3 or Ironman run. A body of research that moved quickly from a niche term to a mainstream one between 2023 and 2026 argues that the ability to hold your numbers late into a long ride — now commonly labelled "durability" — deserves to sit alongside VO2max, exercise economy and fractional utilisation as a genuine, fourth determinant of endurance performance, not just a vague sense of "fading late" (Experimental Physiology, 2025).
For age-groupers, durability is arguably the best available explanation for the most common long-course failure mode: strong watts in testing, a fade in the back half of the bike leg, and a run that comes apart in the first two kilometres. This article walks through what the durability evidence actually says, what a 2025 field study of amateur cyclists found when it looked specifically at age-group athletes, and how you can build a rough version of a durability test into your own training without a lab, a coach, or specialised equipment.
Durability Is Not FTP or Critical Power Wearing a New Name
If you've read TriForward's earlier look at whether FTP is "dead" and whether Critical Power is the better metric, it's worth being explicit about how this topic differs, because the two are easy to blur together. The FTP-versus-Critical-Power debate is about how to test the *ceiling* of your power-duration curve: the ceiling of what you can produce for a given duration when you're fresh, and which testing protocol estimates that ceiling most accurately. Durability asks a different question entirely. It isn't "what's your ceiling," it's "how much of that ceiling do you still have left after hours of accumulated work." You could have a precisely measured, individually calibrated Critical Power and W′ and still have no idea whether your 20-minute power collapses by 8% or 28% after four hours of steady riding. That collapse — not the ceiling itself — is what durability tries to capture, and it has its own separate 2025–26 literature, its own emerging field test, and no dependency on which threshold metric you personally prefer.
The practical reason this matters for triathletes specifically is pacing risk. A rider with excellent fresh numbers but poor durability can look identical to a rider with more moderate fresh numbers but strong durability — right up until hour three, when one of them is still holding goal watts and the other is quietly losing 15–20% of theirs without necessarily feeling like they're "blowing up." The second rider usually still finishes the bike leg on time. They just do it by recruiting more of their remaining muscular and metabolic reserve to hit the same power, and that reserve is exactly what the run then needs.
What the Evidence Says
The Framework: A Proposed Fourth Performance Determinant
The clearest signal that durability has moved from training-forum jargon to a formal research construct is a 2025 methodological paper in Experimental Physiology, which lays out durability as a fourth pillar of endurance performance alongside VO2max, exercise economy and fractional utilisation of VO2max (Experimental Physiology 110(11):1612–1624, 2025). That's a meaningful upgrade in how seriously the idea is being taken. But the same paper is refreshingly candid about its own limits: it states plainly that durability assessment has not been standardised, and that testing duration, intensity, nutritional status and environmental conditions all substantially influence what a durability test actually measures. In other words, the concept is well supported; the measurement of it is still a work in progress. Any protocol you run at home should be read the same way — a useful personal signal, not a certified lab value.
The Amateur Study: A Real Signal, on a Small Sample
The most directly relevant piece of evidence for this article's audience is a 2025 field study in Frontiers in Sports and Active Living, because it is one of the few durability studies run on actual age-group athletes rather than elite or university-lab cyclists. Researchers recruited 14 well-trained amateur road cyclists (mean age 37.5 years, VO2max 52.0 ± 7.4 ml/kg/min, training roughly 9.6 ± 2.2 hours per week) and had them complete a standardised fatiguing protocol at home, using their own equipment (Frontiers in Sports and Active Living 7:1530162, 2025). The athletes were split retrospectively into "successful" and "less successful" groups based on prior-season race results, and the finding was notable: the successful group was distinguished by smaller power declines after the fatiguing protocol, not by higher fresh-state power. Fresh watts looked similar across both groups. What separated them was how well those watts held up.
It's important to be honest about how thin this particular plank is, because it's the one doing the most work in justifying "durability matters for age-groupers" rather than just "durability matters for pros." The study recruited 14 athletes against a planned sample of 26, split into two groups of 7, with "success" defined after the fact from the athletes' own prior results rather than through a controlled intervention. That's a real, amateur-specific data point, and it points in a sensible, mechanistically plausible direction — but n = 14 with retrospective grouping is suggestive, not causal. Treat it as one useful clue, not proof that training durability will improve your results.
Decoupling as a Field-Test Proxy
Running a full fresh-and-fatigued power test isn't the only way durability shows up. A second 2025 paper in the European Journal of Applied Physiology pooled data from 51 trained cyclists across four separate studies and looked at heart-rate decoupling — the drift between heart rate and power output over a long steady ride — as a cheaper proxy for durability that doesn't require a second maximal effort (European Journal of Applied Physiology 125(10):2911–2920, 2025). The correlation between HR decoupling and the actual change in power at the first ventilatory threshold over roughly 2.5 hours of riding was strong (r_rm = −0.76, p < 0.001): more decoupling, more genuine loss of durability. Respiratory-frequency decoupling correlated more weakly (r_rm = −0.40, p = 0.013), and ventilation decoupling didn't reach significance (r_rm = −0.25, p = 0.136). Practically, that means the decoupling number many power meters and head units already calculate on a long ride is a reasonable low-cost early signal of how durable you are, even without a formal two-part fatigue test — though it's a proxy, not a substitute for the real thing.
What's Not Yet Supported
It's worth being equally clear about the claims this research does not yet back. Coaching content on durability, including widely read pieces from practitioners, sometimes asserts that specific session types "build durability best," or that high-intensity interval work can actively *reduce* durability if overdone (CTS/TrainRight, updated 2026). Those are reasonable coaching hypotheses, and they may well turn out to be correct, but they appear in practitioner writing without a cited controlled trial behind them. Present them as experienced coaching opinion, not settled science. Similarly, there is no published triathlon-specific durability intervention trial, and no evidence yet that improving your bike durability specifically improves your subsequent run — a link every triathlete intuitively assumes, and a link nobody has actually tested end to end.
Practical Application: Testing Durability at Home
The genuinely useful part of this research for an age-grouper is that a rough durability check needs nothing more than a power meter, a stretch of road or a trainer, and about three to four hours. The structure mirrors the protocol used in the amateur study above, simplified for home use:
- Warm up properly, then complete a fresh maximal 5-minute effort and a fresh maximal 20-minute effort (on separate days, or with full recovery between them, if you want cleaner numbers).
- On a separate key session, accumulate 2–3+ hours of steady, moderate-intensity riding — the "fatiguing" block. This should feel like solid aerobic work, not an all-out grind.
- At the end of that fatiguing block, repeat the same 5-minute and 20-minute maximal efforts.
- Express each as a percentage decline from your fresh numbers. A smaller decline — power holding up closer to fresh — is the durability signal the research is describing.
- If you'd rather skip the second maximal effort altogether, track HR decoupling across a single long steady ride instead, using whatever your head unit or training platform already calculates; treat a rising decoupling trend over weeks as early warning, not a one-off result.
Because durability assessment isn't standardised, don't over-interpret a single test. Track the trend across a season rather than treating any one number as a verdict on your fitness.
Common Mistakes
- **Turning the test into an unplanned overload week.** A fresh-fatigue-fresh protocol is, by design, a hard day. Squeezed in on top of a normal training week, it becomes an accidental overreach rather than a useful data point.
- **Testing while underfuelled.** The methodological review behind this topic notes that nutritional availability materially changes durability results. A test done on inadequate carbohydrate measures how under-fuelled you were, not how durable you are — and repeated long, low-carbohydrate sessions carry a real low-energy-availability risk, not just a measurement problem.
- **Skipping decoupling context entirely.** Athletes who only ever look at fresh 20-minute power and ignore late-ride drift are, by definition, missing the exact signal this research is about.
- **Chasing session-type claims that aren't yet backed by evidence.** Reorganising your whole training plan around an unproven idea that intervals "hurt" durability is premature; treat that specific claim as opinion, not instruction.
- **Adding long time-in-position rides without a bike fit check.** Building durability generally means building total time in the saddle, and more saddle time is a common trigger for saddle-area, low-back, neck and knee issues if your position isn't sound first.
- **Ignoring overreaching or heat-strain warning signs.** A long fatiguing session, especially indoors, adds real heat-strain risk, and a sudden performance drop paired with elevated resting heart rate, poor sleep or mood change is a signal to back off, not push through.
How to Apply This Week
If durability testing is new to you, don't try to run a full protocol this week. Instead:
- Pick one upcoming long ride and treat it as a durability *observation*, not a full test: note your average power in the final 30–45 minutes versus the first 30–45 minutes, or check the decoupling number your platform already calculates.
- If you want to run the fuller two-effort protocol, schedule it as a single key session with an easy day on either side, book it at least a couple of weeks out from any taper, and fuel it properly — this is not the day to experiment with fasted riding.
- Get a bike fit checked before you start deliberately adding long-ride volume in the name of durability, particularly if you've never had one or it's been more than a season.
- If you have any cardiac risk factors, are new to structured training, or notice new numbness, saddle discomfort or joint pain, treat that as a stop-and-get-assessed signal rather than something to train through.
References
- Durability as an index of endurance exercise performance: methodological considerations — Experimental Physiology 110(11):1612–1624, 2025
- Enhanced durability predicts success in amateur road cycling: evidence of power output declines — Frontiers in Sports and Active Living 7:1530162, 2025
- Durability of the moderate-to-heavy intensity transition can be estimated from decoupling of HR and respiratory frequency — European Journal of Applied Physiology 125(10):2911–2920, 2025
- How to test fatigue resistance and improve durability in cycling — Adam Pulford, CTS/TrainRight (updated May 2026)
- Why durability matters more than FTP in triathlon training — Triathlete.com, 2026
- Durability decoded — a 2025 perspective (research round-up) — Scientific Triathlon, Dec 2025
Frequently asked questions
Is durability the same thing as FTP or Critical Power?
No. FTP and Critical Power describe the ceiling of your power-duration curve when you're fresh. Durability describes how much of that ceiling you still have left after hours of accumulated riding. You can have a precisely tested threshold and still not know how well it holds up late in a long race — that gap is what durability research is trying to measure.
How strong is the evidence that durability matters for age-group triathletes specifically?
Moderate. The concept of durability as a fourth performance determinant has a credible 2025 methodological framework behind it, and a 2025 field study of 14 amateur cyclists found that successful athletes were distinguished by smaller power declines after a fatiguing protocol rather than by higher fresh power. That amateur-specific link is real but rests on a small sample (n = 14 against a planned 26), so treat it as suggestive rather than conclusive.
Do I need a lab to test my own durability?
No. A rough version can be done at home: a fresh maximal 5-minute and 20-minute effort, a separate long fatiguing ride of 2-3+ hours, then the same maximal efforts repeated at the end, expressed as a percentage decline. Alternatively, tracking heart-rate decoupling on long steady rides gives a lower-effort proxy that correlated strongly with durability loss in pooled research on 51 cyclists.
Does high-intensity interval training reduce durability?
That specific claim appears in some coaching content but is not backed by a cited controlled trial in the current research. Treat it as an experienced coaching opinion worth considering, not an evidence-based rule to restructure your whole training plan around.
Will improving my bike durability also improve my triathlon run?
That link is intuitive and widely assumed, but no published triathlon-specific trial has actually tested whether improving bike durability improves the subsequent run. It's a reasonable hypothesis given how durability testing works, not a proven outcome.
How often should I run a full durability test?
Sparingly. Because it's a genuinely fatiguing protocol, treat it as a planned key session with easy days on either side, done no more than a few times per season, and never scheduled during your final taper.
