Claude Opus 5 and your training: what actually gets better, and what does not
A better model improves the parts of coaching that are language and judgement. It does nothing for the parts that are physics and adherence, and those are the parts deciding your result.

Anthropic released Claude Opus 5 on 24 July 2026. Within a day the fitness corner of the app stores did what it always does: "now powered by the world's most advanced AI", new screenshots, same workout. If you train, and you pay a few pounds a month for something with the word "AI" on the icon, you are entitled to ask a blunt question. Does any of this change what happens to your body?
The reasonable assumption is that a smarter model means a smarter programme, and a smarter programme means better results. That assumption is half right, and the half that is wrong is the expensive half. Model quality moves the parts of coaching made of language and judgement: explaining, adapting, answering, substituting, reading what you actually meant. It moves nothing in the parts made of physics and behaviour: the sets still have to be done, the protein still has to be eaten, and nobody has yet built a model that can turn up to the gym on a Tuesday for you.
The more interesting story is not the benchmark table at all. It is the price. What a release like this really decides is whether a genuinely good AI coach can exist inside a normal consumer subscription, or whether decent coaching stays locked behind a service costing as much per month as a gym membership costs per year. That is the part worth 2,000 words, and it is where we will spend most of them.
Train with a programme that can explain itself, with Pocket Fit. Pocket Fit builds your week from your goal, days, equipment and injuries, then progresses it on a written rule rather than a guess. Free on the App Store and Google Play, no card needed.
Our bias, stated first
We build Pocket Fit, which is an AI fitness app. We have an obvious commercial interest in you believing that AI coaching is useful. Weigh everything below against that, and if it ends the conversation for you, that is a fair call.
What we can do is make the piece checkable. Every claim about Claude Opus 5 comes from Anthropic's own announcement, published 24 July 2026, and is linked at the bottom. Every claim about training or behaviour is tied to a named, peer-reviewed paper with a DOI. Where we could not verify something, we left it out rather than filled it in.
One more thing, and it is the most important sentence in this article. We are not going to tell you which model sits behind our in-app assistant, and we are not claiming this one does. Model names change every few months. An app that puts a model name in its marketing is telling you about its supplier, not about your training. Treat "powered by [latest model]" the way you would treat a restaurant advertising the brand of its oven.
What actually shipped on 24 July 2026
The short version, from Anthropic's announcement. Opus 5 is available across Claude.ai, Claude Code, Claude Cowork and the API, where it is called claude-opus-5. It is priced at $5 per million input tokens and $25 per million output tokens, the same as Opus 4.8. A "fast mode" runs at double the base price.
On the benchmarks Anthropic published, it is a large step rather than a polish:
| Benchmark | Reported result |
|---|---|
| Frontier-Bench v0.1 | Surpasses all other models and more than doubles Opus 4.8's performance |
| CursorBench 3.2 | Within 0.5% of Fable 5's peak score, at half the cost per task |
| ARC-AGI 3 | Score three times as high as the next-best model |
| OSWorld 2.0 | Beats Fable 5's best result at just over a third of the cost |
Anthropic also describes it as a thoughtful and proactive model with much stronger visual outputs, and with greatly improved performance on software engineering and scientific research tasks. The announcement did not specify a context window, so we are not going to invent one.
Now hold those numbers loosely, because none of them are about you. ARC-AGI is an abstract reasoning puzzle set. OSWorld is a computer-use benchmark. Neither has ever squatted. The relevance to your training is entirely indirect, and it runs through one channel: cost.
The number that matters is not the benchmark, it is the price
Here is the economics, in plain terms, because almost nobody explains it to the person paying the subscription.
Every question you ask an AI coach costs the company that built it real money, per question. Not a rounding error at scale. Text going in is charged, text coming out is charged, and reasoning models tend to produce a lot of text. That single fact has quietly shaped every AI fitness app you have ever used.
It leaves a developer with three options, and you have met all three:
- Use a cheap, weak model. Answers are fast and shallow. Ask why today is 3x5 and you get a paragraph of encyclopedia. Ask for a substitution and you get a generic swap that ignores the injury you just described. This is most free AI features in fitness apps.
- Use a strong, expensive model and ration it. You get ten messages a month, or a credit meter, or a hard cutoff mid-conversation. The quality is real and you cannot have very much of it.
- Use a strong model and charge properly for it. This is the human-adjacent end of the market. Future, a remote coaching service pairing you with a real human, ran at around $199 per month at the time of writing. That is a defensible price for a person. It is not a price most people will pay for software.
Meanwhile a typical fitness app subscription is a few pounds a month. The gap between those two numbers is the whole problem. We wrote about that spread in detail in our comparison of the best AI personal trainer apps, and the taxonomy there still holds: four unrelated technologies are all sold under the phrase "AI coach", at prices differing by a factor of fifty.
A price drop at constant quality is the only thing that collapses that trade-off. When near-frontier capability costs half of what frontier capability costs, the third option stops being the only way to get good answers, and the first option stops being the only way to stay solvent. Opus 5 holding the same price as Opus 4.8 while roughly doubling its Frontier-Bench score is precisely that shape of change: not "the best model got better", but "the good-enough tier got cheaper per unit of quality".
That is what filters down to you. Not as a feature announcement. As fewer message caps, fewer shallow answers, and fewer apps where the chat tab is obviously a toy bolted to the side.
See what a programme that adapts to your week costs. Pocket Fit builds, progresses and reshuffles your training for a fraction of a human coaching retainer. Free on iOS and Android.
What genuinely gets better for you
Now the useful part. Where does model quality actually show up in a training app? Five places, and they are more specific than "it is smarter".
Explaining, not just prescribing. Almost every app can hand you 3x5 back squat. Very few can tell you why it is 3x5 today and not 4x10, in language pitched at your actual level, and then re-pitch it when you say you did not follow. That second part is pure language capability. A weaker model produces a textbook paragraph. A stronger one produces the paragraph you needed. This matters more than it sounds, because understanding why you are doing something is one of the load-bearing behaviour change techniques in the literature rather than a nice-to-have.
Substitutions with judgement. "The squat rack is taken and my knee has been cranky since Tuesday" is not a lookup. It has a constraint, a context and a safety dimension, and the correct answer depends on weighing all three. Weak models pattern-match on the word "squat" and hand you a leg press. Stronger models notice the knee, notice that you asked at 6pm on a busy weeknight, and give you something you can actually load today. This is the single most visible quality difference in day-to-day use.
Reading intent out of messy input. Bringing your own programme in from a coach's PDF, a screenshot, or three lines of shorthand in your notes app is a parsing problem full of human ambiguity. "DB inc press 3x8-10 @RPE8" needs to become structured sets your app can progress. So does "legs, then whatever the shoulder feels like". Pocket Fit already supports importing your own programme and filling the gaps from your goal and experience, and this is exactly the class of task where better models are noticeably better.
Visual understanding. Anthropic describes this generation as much stronger visually. For fitness software that points at two things: food photos and form. Pocket Fit already does photo and label based food logging, and it is fair to say better vision helps the whole category here. It is not fair to promise your dinner will suddenly be measured to the gram, and we are not going to. Photo estimation is still estimation.
Fewer confidently wrong answers. Better models hallucinate less. They do not hallucinate zero. That distinction is the whole reason for the next section.
Why the weight on your bar should never come from a language model
This is the part we would argue for even if it cost us a marketing line.
Anything with a number in it that you are going to act on should come from a deterministic rule, not from a model. Your next working weight is the clearest case. It is arithmetic on your own logged history, and arithmetic has a right answer that does not vary with phrasing, mood or temperature setting.
In Pocket Fit that rule is written down and does not move. Reps go up first, one at a time, to a cap of 15. Weight only rises after a confirmed plateau across two matching sessions, and barbell increments are 5 kg for men and 2.5 kg for women. You can audit it. You can predict it. Run the same history through it twice and you get the same answer twice. We explained the reasoning in full in reps first, then weight.
A language model, however good, gives you a plausible number. Plausible is a lower standard than correct, and for a load you are about to put on your spine, the gap between those two words is the entire safety margin. The correct division of labour is boring and stable: the deterministic engine owns the numbers, the model owns the words. Progression, load, increments, volume. Rules. Explanation, substitution, parsing, conversation. Model.
Any app that blurs that line is making a design mistake, and a new model release does not fix it. If anything it makes the temptation worse, because a better model produces more convincing wrong numbers.
What does not get better at all
Now the contrarian half, and it deserves as much space as the optimistic half.
No model increases your training frequency. Adherence is the variable that decides outcomes, and it is not a software capability. Schoenfeld, Ogborn and Krieger pooled 10 studies for Sports Medicine and found that training a muscle group twice a week produced greater hypertrophy than once a week, at matched volume. The lesson is about sessions done, not sessions planned. A model that writes a better Tuesday is worth nothing if you do not train on Tuesday.
Protein and calories are physiology. Morton and colleagues pooled 49 studies covering 1,863 participants for the British Journal of Sports Medicine and found protein supplementation produced meaningfully greater gains in lean mass and strength, with benefits levelling off at roughly 1.6 g per kg of bodyweight per day. Longland and colleagues ran a randomised trial in the American Journal of Clinical Nutrition in which participants in a steep energy deficit ate either 1.2 or 2.4 g/kg of protein while training hard; the high-protein group gained 1.2 kg of lean mass while the low-protein group essentially held steady. No language model changes those numbers. They were true before Opus 5 and they will be true after Opus 6.
Habits take the time they take. Lally and colleagues tracked 96 people forming new daily habits for European Journal of Social Psychology and found automaticity took a median of 66 days, with a spread from 18 to 254. That is a property of people, not of software.
A model cannot see the room. It cannot watch your hip shift under a heavy set, cannot palpate a sore shoulder, and cannot tell an irritated tendon from something that needs a scan. This is education, not medical advice. If something hurts in a way that persists, changes at night, or gets worse under load, see a physiotherapist or doctor. That advice does not get cheaper or better with model releases.
The bottleneck is almost never the software. For most readers of this article, the difference between a good year and a wasted one is showing up three or four times a week for fifty weeks. Nothing on Anthropic's benchmark table addresses that.
Where the evidence runs out
Being honest about the limits is the point of publishing this rather than a press release.
Benchmark scores are not user outcomes. Frontier-Bench, CursorBench, ARC-AGI and OSWorld measure abstract reasoning, coding and computer use. There is no established measurement linking any of them to whether the person holding the phone gets stronger. The chain from "better on ARC-AGI" to "better squat in October" has at least four unvalidated links in it.
Nobody has run the trial. As far as we can find, there is no randomised controlled trial comparing training outcomes between users of a stronger and a weaker language model inside the same fitness app, holding everything else constant. That study would be straightforward to design and expensive to run, and until someone runs it, everything anyone says about model quality and training results, including everything in this article, is reasoning from mechanism rather than from evidence.
The effect may well be small. Michie and colleagues' meta-regression in Health Psychology, covering 122 physical activity and healthy eating interventions, found that the techniques with the clearest effects were self-monitoring combined with things like goal setting and feedback. Those are structural features of an app, not model features. A deterministic tracker with a plain interface delivers most of them. It is entirely plausible that the marginal effect of a better conversational model on real-world adherence is close to zero.
The behavioural evidence is also mixed on the layer everyone reaches for. Mazeas and colleagues meta-analysed 15 randomised trials for the Journal of Medical Internet Research and found gamification increased physical activity by roughly 1,600 steps per day, but noted the effect on longer interventions was much less certain. We build week-based streaks with earned grace weeks for exactly that reason, and we wrote about why daily streaks are hostile to people who train three or four times a week in accountability without shame.
And finally: we are two days out from the release. Any confident claim about what this model does in production over a year is not available to anyone yet, including us.
What to actually do this week
If you take one practical thing from a model launch, take this.
- Judge an app by its rules, not its supplier. Ask what happens to your weights when you hit the top of a rep range. If the answer is a written, predictable rule, that is a real product. If the answer is a model name, keep looking.
- Use the chat for judgement, not arithmetic. Substitutions, explanations, "my knee is cranky, what should I do with leg day". Not "what should I squat on Monday".
- Fix the physiology before the software. Protein near 1.6 g per kg, calories roughly where they need to be for your goal, sleep on a consistent schedule. Those move the needle more than any release.
- Count sessions, not features. Four sessions a week on a mediocre programme beats two on a perfect one, every time.
Put it together
Model releases are real progress on a narrow axis. Language and judgement genuinely improve, and price per unit of quality genuinely falls, and that second one is what decides whether good coaching lives inside a subscription you can afford or a retainer you cannot. That is worth something and we are not going to pretend otherwise.
But your result is still decided by two things, and neither is on the benchmark table. The first is whether you show up. The second is whether the numbers you act on come from an auditable rule rather than a plausible guess. Pocket Fit is built around that split on purpose: programme generation from your goal, days per week, difficulty, equipment, split, and your free-text limitations and injuries; seven splits with real day constraints; a progression engine that raises reps to a cap of 15 before it ever touches the bar; your own imported programme with the gaps filled from your goal and experience; photo and label based food logging; and week-based streaks with earned grace weeks so a bad week does not cost you the habit. The programme updates week to week from what you log.
Georgi built it because of the bad weeks, not the good ones. He went from 122 kg to competing at The Yard Games, losing 38 kg along the way, and the app exists because the weeks where everything slipped were the ones that decided the outcome. You can read the whole thing in our story. No model release changes that lesson. It just makes the explaining better.
Start with Pocket Fit, free. A programme built round your equipment and your injuries, progressed on a rule you can read. Personalised in minutes on iOS and Android.
Claude Opus 5 and AI fitness apps: common questions
Does Claude Opus 5 make AI fitness apps better?
It improves the parts of an app made of language and judgement: explaining why a session is structured the way it is, substituting an exercise when equipment is taken or a joint is sore, and parsing a programme you paste in from a PDF or screenshot. It does not improve progression maths, which should come from a deterministic rule, and it does not improve adherence, which decides most results.
What does Claude Opus 5 cost?
Anthropic lists it at $5 per million input tokens and $25 per million output tokens, the same pricing as Opus 4.8, with a fast mode at double the base price. That matters to you indirectly: better quality at an unchanged price is what lets an app give you real answers without a message cap or a $199 per month retainer.
Does AI coaching actually work?
The honest answer is that nobody has run the trial. There is no randomised controlled trial comparing training outcomes between a stronger and a weaker model inside the same app. What is well evidenced is the layer underneath: Michie and colleagues' meta-regression of 122 interventions found self-monitoring plus goal setting and feedback were the most effective techniques, and those are structural features rather than model features.
Should an AI pick my working weight?
No. Your next weight is arithmetic on your own logged history, and it should come from a written, auditable rule. Pocket Fit raises reps first to a cap of 15 and only raises weight after a confirmed plateau across two matching sessions, with barbell increments of 5 kg for men and 2.5 kg for women. A language model produces a plausible number, and plausible is a lower standard than correct.
What is the best AI fitness app in 2026?
The useful test is not which model an app uses, it is whether the app has a progression rule you can read, handles your actual equipment and injuries, and does not punish you for training four days a week instead of seven. Model names change every few months. A written progression rule does not.
References
- Anthropic (2026). Claude Opus 5. Announcement, 24 July 2026. https://www.anthropic.com/news/claude-opus-5
- Schoenfeld BJ, Ogborn D, Krieger JW (2016). Effects of Resistance Training Frequency on Measures of Muscle Hypertrophy: A Systematic Review and Meta-Analysis. Sports Medicine, 46(11), 1689-1697. DOI: 10.1007/s40279-016-0543-8
- Morton RW, Murphy KT, McKellar SR, Schoenfeld BJ, Henselmans M, Helms E, Aragon AA, Devries MC, Banfield L, Krieger JW, Phillips SM (2018). A systematic review, meta-analysis and meta-regression of the effect of protein supplementation on resistance training-induced gains in muscle mass and strength in healthy adults. British Journal of Sports Medicine, 52(6), 376-384. DOI: 10.1136/bjsports-2017-097608
- Longland TM, Oikawa SY, Mitchell CJ, Devries MC, Phillips SM (2016). Higher compared with lower dietary protein during an energy deficit combined with intense exercise promotes greater lean mass gain and fat mass loss: a randomized trial. American Journal of Clinical Nutrition, 103(3), 738-746. DOI: 10.3945/ajcn.115.119339
- Lally P, van Jaarsveld CHM, Potts HWW, Wardle J (2010). How are habits formed: Modelling habit formation in the real world. European Journal of Social Psychology, 40(6), 998-1009. DOI: 10.1002/ejsp.674
- Michie S, Abraham C, Whittington C, McAteer J, Gupta S (2009). Effective techniques in healthy eating and physical activity interventions: a meta-regression. Health Psychology, 28(6), 690-701. DOI: 10.1037/a0016136
- Mazeas A, Duclos M, Pereira B, Chalabaev A (2022). Evaluating the Effectiveness of Gamification on Physical Activity: Systematic Review and Meta-analysis of Randomized Controlled Trials. Journal of Medical Internet Research, 24(1), e26779. DOI: 10.2196/26779
Pocket Fit is a fitness and wellbeing app, not a medical device. It does not diagnose, treat or prevent any condition. Always consult a qualified healthcare professional before starting or changing a training or nutrition programme, and if you have persistent problems with sleep, pain or fatigue.
Georgi, founder of Pocket Fit. He went from 122 kg to competing at The Yard Games, having lost 38 kg along the way.
Train smarter with Pocket Fit
Download the app