Writing

AI Fitness Coaching Is a Context Problem

I built my own coach out of three MCP servers wired into Claude. The interesting engineering wasn't the prompt, it was the data.

22 min read

  • AI
  • MCP
  • Developer Tooling

For 18 months I trained hard and ate well. Tonal for strength, rucking for everything else, and a diet I actually stuck to. The effort was never the variable.

The variable was knowing what to do on any given day, because the choice between going heavy, backing off, or skipping entirely is a real decision with consequences. Get it wrong in one direction and you leave progress on the table. Get it wrong in the other and you are overtraining, going backwards, and risking injury.

Here is what I had to make that call with. My Garmin CIRQA knew how I slept, what my HRV and resting heart rate were doing, and how much I had moved. My food log knew what I had eaten. My Tonal knew every set I had lifted and at what weight.

Tonal strength score beside the Garmin Connect home screend

Two apps, two truths, no overlap. Tonal knows my strength training. Garmin Connect knows how I slept. Neither one knows the other exists.

Neither one could see the other, which meant both were giving me reasonable answers to incomplete questions. Tonal would green light a heavy session without knowing how I had slept, what my heart rate had been doing overnight, or that I had rucked seven miles the day before. Garmin would hand me a training readiness score with no idea what I had actually lifted or how much I had eaten.

I was already asking Claude about all of this, and the chats were superficial for an obvious reason: I would screenshot the Garmin sleep screen, screenshot the Tonal summary, paste them both in, and type out what I had eaten from memory. It worked, in the sense that I got answers, but I was spending more time assembling the context than acting on it. That is usually the point where you stop doing a thing by hand.

Any decent model already knows more about progressive overload, protein targets and recovery than I ever will. What no model has is what you did yesterday.

The plumbing was missing, not the model.

The result, so you know whether this was worth doing: these are my personal results, and this tool helped me get there, it did not take away the effort. Over those 18 months I landed under 15% body fat, down roughly 44 pounds of fat and up about 9 pounds of lean mass. Adding muscle through a deficit that long is the part I actually care about.


What I actually wanted

I wrote these down before building anything, and every decision later in this piece traces back to one of them.

  1. One conversation, all three streams. Not a dashboard, and not three integrations I still have to reconcile in my head. One place where I can ask a question that touches sleep, food and lifting at the same time.

  2. It has to work from my phone. I need my coach with me, in the kitchen and at the machine and out on the trail, rather than waiting at my desk for me to come ask it something.

  3. I own my trend data. I am not pretending to be off SaaS here: Tonal and Garmin Connect are both cloud services and both of them own the raw feed. What I wanted was a local copy of the history that matters, so that the eighteen months of weight, intake and session data behind every decision is mine. Hardware gets replaced and platforms get abandoned, and when I eventually move off one of these I want to carry the trend with me rather than start again at zero.

  4. It has to write, not just read. A system that can tell me what to do but cannot log my weight or build tomorrow’s workout is a very expensive dashboard.

  5. The loop has to close. Garmin is my system of record for everything my body does, so anything that happens outside it has to find its way back in. If a Tonal session never lands in Garmin, then Garmin’s view of my week is wrong, and every recommendation built on that view inherits the error.


What it looks like now

Morning

Logging weight and blood pressure, then asking whether I recovered enough to train hard

I step off the scale and say “log 190.2.” It goes to Garmin, and what comes back is where that sits against last week rather than a number in isolation. Blood pressure from a Microlife WatchBP Home A goes in the same way.

Then I ask whether I recovered enough to go hard, and it reads last night’s sleep architecture rather than just the total: 8 hours 7 minutes, sleep score 90, 115 minutes of deep and 115 of REM, only 9 minutes awake, resting heart rate 55. Alongside that it reads overnight HRV, which Garmin derives from beat-to-beat variation during sleep (rMSSD, measured in milliseconds) and reports as a seven-night average against a three-week baseline. A single night of poor HRV means very little on its own, but a seven-night average drifting below baseline while resting heart rate climbs is the pattern worth acting on, and that combination is exactly what I want something else watching for me.

That morning it told me this was my best sleep of the block by a clear margin, and that nothing in the numbers said hold back.

Eating

Storing a food in the library, then recalculating the day against my macro targets

Logging has to be nearly free or it does not happen. Most of the time I am not photographing anything, I am typing “log choc p-shake” or “log ground beef + tomato sauce,” because those are already saved and it knows what they are. When I do want precision on something unfamiliar, a photo of the plate works.

Later I ask what is left, and on a Friday training day it tells me I have used 574 calories and 83 grams of protein against a target of 1,288 and 190, leaving 107 grams of protein inside 714 calories.

The useful part is what happens next, because it does not hand me a macro puzzle to solve. It suggests actual meals, drawn from the foods I have saved, the recipes I have built, and what I have logged on similar days before. That is the difference between a calculator and something worth talking to.

Training and problem solving

Claude editing Monday's push workout in Tonal and reporting what changed

Claude's edited Monday push workout in Tonal, with Standing Incline Press in place of the flys

Monday’s push workout, corrected: flys dropped, incline pressing in their place, and I never opened the app to make the change.

This is the part that changed my training.

I had been failing partway through sets on pressing movements, with a pinch in my back that came along with it. We worked through the pattern over several weeks, and the actual cause turned out to be bigger than one bad exercise: I was training chest two days in a row, Sunday into Monday, without enough recovery between them. That is not something you catch from a single rough session. It only surfaced once we were watching recovery signals and progress over several weeks rather than any one workout in isolation. Once we saw it, the fix was two changes: move Sunday’s chest work onto Monday so it stopped repeating two days running, and on what was left of the Monday session, drop the flys and switch to incline pressing. Claude made those edits directly in Tonal through the MCP server, so the corrected session was simply waiting for me the next time I walked up to the machine. I never opened the Tonal app to rebuild anything by hand, which is a huge time saver.

The following Monday I did the session and reported back that it went fine, not especially difficult.

The response was that this was the point. Average heart rate of 100, barely into zone 2, no struggling at the end of any set, and a clean session with reps left in reserve. Given how the previous weeks had gone, an unremarkable session was the goal, because it meant the two changes had put me in a range I could complete properly rather than one sitting right at the edge of what I could manage.

I would have read that same session as a wasted workout, an easy day when I had wanted a hard one. The system read it as evidence the fix had worked, and it could only reach that conclusion because it had the failed sets from three weeks earlier, the change we made in response, and the heart rate data from the session itself, all available at once.

None of that required a smarter model, just the last month of my training history somewhere it could actually reach.

Pulling activity data

Pulling the day's ruck from Garmin and reading the zone distribution against the calorie estimate

Anything Garmin has recorded is just as easy to pull back out. “Did you pull my ruck today” returns the full activity detail: pace, heart rate zones, training effect, and a Pandolf-corrected calorie estimate that runs meaningfully lower than Garmin’s own number for the same session.

Evening

Logging the evening's food and checking what's left against the macro target

At the end of the day I get a recap: what I actually ate against the target, what the session did to my weekly load, and what tomorrow looks like given both. If Friday came up short on protein, Saturday’s plan already accounts for it. The waist measurement gets checked weekly rather than daily, because that is the number that tracks with real progress while the scale is mostly noise.


Everything above is the story. Everything below is the build. If you are here for the systems, this is where it starts.


The architecture

Four components, running in Docker on a Windows box in my house that never sleeps.

Claude and the skills feeding through the tunnel into three MCP servers, with the sync loop running back into Garmin

garmin-mcp is the read path for everything my body does without being asked. It is FastMCP and python-garminconnect over a local SQLite cache, exposing about a dozen read tools covering readiness, sleep, daily stats, training load, activity detail, body trend, zone summary and VO2 max, plus a small set of writes for weight, blood pressure and hydration, each with a matching delete.

macro-mcp handles nutrition, and it is the component I underestimated most. It covers food logging, a personal food library, recipes, body composition, progress photos, targets that vary by day type, and trend queries. My Tuesday target is 1,120 calories on a ruck day, Sunday is 1,930 on a refeed, and Friday is 1,288 on a training day. Those are stored per date rather than derived, for a reason I will come back to.

tonal-mcp is the write path into training. It searches Tonal’s movement library, creates workouts, updates them when the plan changes, and estimates how long a session will take. This is what makes the system a coach rather than an advisor, because when we decide to drop the flys and switch to incline, that edit lands in the machine I am about to stand in front of.

tonal-garmin-sync is not an MCP server at all, and it is the component that makes the rest of the system trustworthy. It is Node and TypeScript in Docker: a webhook server triggered by Home Assistant, a wrapper around Tonal’s API, a FIT file encoder, a movement mapper covering 291 Tonal exercises, a dedup store, and a Python uploader. A finished workout becomes per-set detail, then a strength FIT file, then a real activity in Garmin Connect with every set, rep, weight and heart rate sample intact.

This one is a fork, not something I built from scratch. The core sync engine, the webhook trigger, FIT encoding, the duplicate-free upload, the original 291-exercise movement map, weight-doubling for both-arms movements, ruck load parsing, weather capture, and the Tonal calorie-fallback logic, all came from theengineer1676/tonal-garmin-sync. What I added on top is workout genre classification, so a yoga or mobility session syncs to Garmin with the correct sport type instead of getting logged as strength.

The Python sync job pulling weight, RHR, profile and FIT data from Garmin Connect into the local cache

The job that keeps that cache warm: a partial pull of the last hour’s data every 20 minutes, plus a full refresh overnight in case an older activity gets backfilled.

Tonal's Wednesday lower session next to the same workout's set-by-set detail inside Garmin Connect

The sets match, rep for rep. What Tonal logged on the left is what shows up in Garmin on the right, same weights, same reps. The exercise names occasionally don’t, since Garmin’s library is a subset of Tonal’s and some mapping loss is unavoidable.

The arrow pointing back into Garmin is the part people miss when they look at this diagram. Everything else in the system reads Garmin’s training load, so if strength work never lands there, the whole thing is reasoning confidently from a number that is missing half my week. Fixing that input mattered more than anything I did on the model side.


Three layers, not one

The data plane is the MCP tools, which are deterministic, boring and opinion-free. They return numbers with units and hold no view about what those numbers mean.

The semantic layer is the skills. garmin-coach and macro-coach teach the model which tool answers which question, what each field actually measures, what the units are, and what a null means. That is the entire scope.

Judgment is the model reasoning against the plan I have described in conversation. Worth noting that I run almost all of this on Sonnet rather than Opus. For day-to-day logging and readiness checks the answers were not better on the larger model, and it was noticeably slower for what amounts to a series of small, well-scoped tool calls.

The rule I enforced throughout: the skills contain zero training philosophy. No thresholds, no “if HRV drops below X then deload,” and no programming rules of any kind.

That looked like an omission when I wrote it, and it is the reason the system still works. A training approach changes every few weeks, while tool semantics do not, so coupling them means every change to my program becomes a code change and a redeploy. Worse, a hard-coded threshold is a coach you cannot argue with, and the entire value of this thing is that I can say “I am three months into a cut and traveling next week” and have that shift the answer.

The skill’s job stops at the API. Anything it might guess about me instead belongs in a conversation.

If you have ever pulled policy out of platform code so the platform stops needing a release every time policy changes, this is the same move.


Decisions and tradeoffs

Three servers instead of one. They fail independently, they iterate independently, and I can enable only the ones a given conversation needs, so a nutrition question does not have to carry the Tonal tool list. A single combined server would put every tool in front of the model on every turn whether or not the question is about lifting. Three deploys instead of one is a real cost, but it’s an operational one, not an architectural mistake.

Streamable HTTP behind a Tailscale Funnel. The phone requirement meant the servers had to be reachable from outside the house, and this is the cleanest way I found to do it. No open ports on the router, no certificates to renew, and no reverse proxy to babysit, because Tailscale terminates TLS at its edge and the machine itself never sits directly on the public internet.

Two undocumented vendor APIs. Garmin’s and Tonal’s, with no SLA, no contract, and no warning when they change. Tonal has no OAuth for this, so the sync stores a password in plaintext, which I would not accept anywhere else. It is a single-user system on hardware I control, and that is the only reason it is acceptable here.

A local SQLite cache. A 90-day trend question should not turn into 90 API calls.

Trend tools instead of raw dumps. get_body_trend, get_intake_trend and get_activity_trend do the math on the server and return the answer. If I ask for ninety days of raw entries instead, the model has to read all of it just to work out a direction, and there is less room left for the actual question. Let the server compute the slope and send the slope.

Write tools, including deletes. These are the ones that can actually hurt you. A model will eventually mis-log something, and a coach that cannot correct its own mistake is one you stop trusting, so every write has a matching delete. delete_workout goes a step further and returns a snapshot of what it just removed, because Tonal strips the sets off an archived workout and an archived workout that cannot recreate itself is not much of a safety net.

Two test suites, because mocked tests were not enough. The fast suite is mocked and runs offline: 134 tests in macro-mcp, 83 in tonal-mcp, covering request shapes, conversions and error handling. It runs on every change.

But mocked tests are only as correct as the fixtures you hand them. A wrong assumption about a third-party API’s shape produces a test that passes for the wrong reason, which is how a bug requiring weightPercentage to be an integer got through: the tests called the service layer directly and skipped FastMCP’s schema coercion, so the part that was actually broken was never exercised.

So there is a second suite, marked @pytest.mark.integration and excluded from the default run. It needs real credentials, goes through the actual tool-call path, and creates and archives real throwaway workouts in my account. The rule it enforces is narrow and useful: every behavioral claim in a tool’s docstring gets a live assertion, not just a call. If the docstring says an archived workout stays listed but loses its sets, a test round-trips exactly that and checks both halves.

The first version was a standalone script, and when an assertion failed partway through it died before reaching cleanup and left an orphaned test workout behind in my account. The pytest version uses a fixture with try/finally teardown, so a failing test still archives whatever it created.

Null means “no data yet,” not “not built.” This actually happened. I had not set targets for a date, the tool correctly returned null, and the model told me the feature had not been implemented yet. The tool was right and the answer was wrong, so the fix went in the skill rather than the code: say plainly what an empty result means. When your caller is a language model, that documentation is part of shipping the tool.

A personal food library. Logging the same shake three hundred times should not cost three hundred estimations. I tell it once what a Fairlife is, and after that “log fairlife” is the whole interaction. The deeper reason is that consistency beats precision in a food log: a slightly wrong number applied identically every day still produces the right trend, while a freshly estimated number every day adds noise to the only signal that actually matters.

estimate_workout_duration as a tool. Models are bad at guessing how long a session will take, so give it a calculator rather than asking it to reason.

A hand-maintained 291-movement map. Tonal has movements that Garmin’s FIT format simply has no name for, so every Tonal exercise has to be mapped onto the closest thing Garmin understands, and I maintain that mapping by hand. It is approximate by nature and always will be. Where there is no good match the exercise still syncs, it just arrives labelled “Unknown” with the weights and reps intact. Rejecting the workout entirely would trade real set data for a cosmetic label, so I’d rather it degrade partially than fail closed.

Event-driven where possible, polling where not. Home Assistant fires the webhook on Android, while iOS falls back to polling. Sometimes your architecture is decided by someone else’s notification model.

Looking at that list again, most of these are not really about the model at all. The cache and the trend tools are about how much room it has to think, null semantics and the movement map are about whether the data means what it says, and the undocumented APIs are about whether the data shows up in the first place.

Where it is still rough

  1. Authentication. Tailscale Funnel gets traffic to the house safely, but there is no real authorization layer in front of the servers themselves. That is fine for one user on hardware I own, and it is the first thing you would need to fix for anyone else.

  2. No approval gate on writes. The model can log a weight or edit a workout without checking with me first. It has not caused a problem yet, and “yet” is doing some work in that sentence.

  3. Skills inform, they cannot enforce. The skill can tell the model that a relative date is ambiguous and that it should confirm before writing, but it cannot make it do so. If I say “log yesterday’s ruck” at one in the morning, nothing in the system stops it from picking the wrong day. Guidance makes a mistake unlikely without making it impossible, and that distinction matters once a system starts writing rather than just reading.

  4. Running the integration suite makes me slightly nervous. It authenticates as me and writes to a live account through an API nobody documented or licensed for this, and I have not read anything in the terms of service that clearly says it is fine. So I have run the full suite only a handful of times, enough to confirm the writes round-trip correctly, and I lean on the mocked suite for everything day to day. That is a real constraint rather than a preference: the tests I most want to run continuously are the ones I am least comfortable running continuously.

  5. It runs on a Windows box in my house, which is obviously not ideal. If that machine reboots and something does not come back up, the coach is offline until I happen to notice. Moving the containers onto ECS or any small managed host would fix it and would not be hard, but for a single-user system whose worst case is logging a weight an hour late, it has not been worth doing yet.

  6. Single user by design. Multi-tenant is a different program rather than a bigger version of this one.

  7. The undocumented APIs are a ticking time bomb. Both of them will break eventually and I will not get a deprecation notice when they do. The local cache is the mitigation, since the coach still has ninety days of history to reason from when an API goes down, which means it degrades rather than dies. That is about the most you can do about a dependency you do not control.


Running it yourself

Four repos, and I would do them in this order, because each one is useful on its own and the later ones assume the earlier ones already work.

1. garmin-mcp (README) Start here, and start read-only. Get get_daily_stats returning before you do anything else. You need a Garmin account and Docker, and this component alone will already tell you more about your own training than the Garmin app does.

2. Add the tunnel and test it from your phone. Tailscale Funnel, or Cloudflare Tunnel if you prefer. Do this before building anything else, because if remote access does not work then everything downstream is wasted effort and you want to know that on day one.

3. garmin-coach (README) garmin-coach is a single SKILL.md that teaches Claude how to read what the server returns: which tool answers which question, what the units are, which numbers are comparable to each other, and what a null actually means. It is impersonal, it is reusable by anyone with a Garmin, and it contains no opinion about training whatsoever. Install it through Claude’s Skills UI, or drop SKILL.md into your skills directory depending on how you are running Claude. Then ask the same question you asked in step 1 and compare the two answers.

This is the step people skip and it is the one that matters most, because you add zero new tools and the answers get materially better. There is an evals/evals.json in the repo if you modify the skill and want to confirm it still triggers and reasons the way you expect.

Everything it refuses to hold lives in the conversation instead. My weekly split. My macro targets by day type. The correction factors I apply to Tonal’s calorie estimate and the Pandolf equation I use for rucks. The accumulated findings, like the fact that a bench warm-up ladder is the difference between a good session and a backwards one, or that reading HRV and resting heart rate together is meaningful while either one alone usually is not.

A redacted version of that context is here. Personal figures are placeholders. Swap them for your own and it becomes your coach rather than mine.

4. macro-mcp (README) Nutrition, targets and body composition. It is independent of the training side, so it runs happily with or without the rest.

5. tonal-mcp (README) and tonal-garmin-sync (README) These two only help directly if you own a Tonal, which I realise is niche. The pattern is not niche at all, though: what they do is read a structured strength log and write structured workouts back, and any tracker with an API can fill that role. wger is the obvious substitute, being self-hosted, open source, REST API driven, and covering both workouts and nutrition. Swap it in behind the same tool surface and the rest of the system does not notice the difference.

Each README carries the tool list, the environment variables, a compose file, and an honest warning that all of this is built on APIs nobody promised me.


What it actually took to build

The model was the easy part. Swap it for whatever is best next year and the system still works, because almost nothing I built depends on which model is answering.

Writing the code was not the hard part either. The stack came together quickly. What took roughly two months was the tuning: running it on myself every day, noticing where an answer was confidently wrong, and tracing that back to a tool that returned the right number with the wrong meaning attached. None of that work looks like much in a commit history. It is the difference between a system that technically functions and one you actually trust at six in the morning.

What data actually moves the needle

Almost all of those two months went into working out which numbers changed a decision and which ones merely looked interesting on a chart.

Several metrics I was certain I needed turned out to be things I would look at, nod at, and then do nothing about. Daily step count was the obvious one: it moved with whether I had rucked, which I already knew, and no step total ever changed what I did next. Garmin’s stress score was the same, interesting to look at and never actionable. VO2 max moves so slowly that checking it more than monthly is theatre. And “calories burned” per session turned out to be the most misleading number of the lot, precise-looking and wrong enough that I stopped reading it and started applying my own correction instead.

What survived is a short list:

  • Seven-night HRV average against baseline. Single nights are noise. The trend is not, and of everything on this list it is the single biggest readiness signal I have. It is the clearest read on how much stress my autonomic nervous system is still carrying from training, heat, bad sleep, or just a hard week, and it tends to move before I consciously feel run down, not after. When the trend and how I feel disagree, I have learned to trust the trend.
  • Resting heart rate, read alongside it. Both drifting the wrong way at once is the signal. Either one alone usually is not.
  • Waist, measured weekly. The number that actually tracks with progress while the scale is mostly water.
  • Whether I hit protein. The one nutrition input that reliably predicts whether muscle survives a deficit.
  • Top-set performance over per-set weight. A session’s early sets are a ramp, not the test. On a lower day I might work Romanian deadlifts up through 100, 138 and 146 lb before a top set of 155 lb for 8 reps, and only that last number is worth tracking, because the sets before it exist to get me to a weight worth measuring, not to be measured themselves. Auto-regulating means that ramp shifts around depending on how the day feels, so watching every set mixes signal with noise. Track the top set across weeks and you’re looking at strength. Track every set and you’re mostly looking at how warmed up I felt that day.

Five numbers. Every one of them changes something I do.

That generalises well beyond fitness. Data that never changes a decision is just noise with a timestamp on it. Collecting more of it was never the point. Getting the few numbers that genuinely matter into one place, at the moment the decision actually gets made, was.

And the thing that finally made an AI coach useful was not making it smarter. It was making sure that when it said “go heavy today,” it knew what I did on Tuesday.