I built some tools recently to understand my own fitness, using my Apple Watch data. I learned what I was after about my training, and something I wasn’t looking for about dashboards. Two findings changed how I read every number a machine hands me.
Two numbers I stopped trusting
The first was my VO2 max, and it read artificially high. Sadly, I’m not an Olympic-caliber athlete. So I started wearing a heart monitor strap and compared it against the watch, and that’s how I found the problem: the optical sensor built into my Apple Watch misses brief spikes, the kind you get running or biking up a short hill. Drop the hardest few seconds from the record and the math concludes you produce a given pace at a lower heart rate than you actually do. This is an n of 1, my workouts and my watch, so take it as one person’s observation rather than a verdict on the hardware. The pattern was consistent enough, though, that I stopped taking the number at face value.
The second was stranger. I pointed an AI at my health data and it told me I wasn’t getting enough sunlight over the winter. Alarming, except in my case it’s almost certainly false. The daylight estimate comes from the ambient light sensor on the watch, and from late fall to late spring I’m outside with a long sleeved shirt or jacket over it. If the same estimate leaned on GPS instead, it would tell the opposite story: plenty of time outdoors. The number wasn’t measuring whether I walk the dog enough. It was measuring how dark it was inside my sleeve.
Two sensor design decisions. Two wrong numbers. Neither decision is documented anywhere I’ve seen: nothing warns you that the daylight estimate falls apart once it’s cold enough for long sleeves, or that the VO2 max number rests on assumptions your workout might not match. If I hadn’t dug in, I’d feel great about my fitness and guilty about my sunlight exposure, and I’d be wrong on both counts. The dog would have enjoyed the extra walks, at least until I noticed they didn’t change the number.
The presentation got better. The measurement didn’t.
Garbage in, garbage out is as old as computing. What’s new is how convincing the output looks. A number from a fancy device, on a polished dashboard, with a trend line and a health ring around it, reads as truth. Feed it to an AI and you get confidently delivered, but misleading advice, with the sensor’s blind spot laundered out somewhere along the way.
The model strips the provenance off a wrong number, so a measurement artifact comes back as confident advice with no trace of how it was measured. The sleeve gets scrubbed out. What reaches you is “you need more winter sun,” clean and directive, and nothing in the sentence reminds you that the guidance came from a sensor sitting under a sleeve.
That laundering is the dangerous part, and the reason is that a number and a sentence invite different kinds of doubt. A number has a magnitude, so you can compare it to what you expected and find it too high or too low. That’s what made me look at the VO2 max in the first place. Advice has no degree. You can take it or leave it, but questioning it means reconstructing where it came from, and that’s work you have to choose to do at the exact moment the analysis is telling you it’s already settled.
And you only go reconstructing what bothers you. I checked both of these because both bothered me: the VO2 max was flattering past the point of plausibility, and the sunlight warning accused me of something. If the same broken sensor had told me I was getting plenty of winter light, I’d have believed it and moved on. The wrong answers that survive are the comfortable ones.
Every company has a sleeve
Companies are wiring agents to their data right now, and the corporate version of my sleeve is everywhere. The CRM field nobody updates. The survey only the happy customers bother to answer. The metric that quietly changed definitions last year and kept its name. The agent won’t know any of that. It will analyze what it’s given, confidently, and hand back a clean conclusion with the provenance washed off, exactly the way the sunlight warning came back to me.
The fix isn’t to measure less. It’s to treat “how is this measured?” as part of reading any number, the same way you’d read the byline on an article. And it’s to stop assuming that some measurement always beats none. A wrong number you trust does more damage than a gap you know about, because the gap keeps you looking and the wrong number makes you think you can stop.
None of this is new, exactly. A spreadsheet could always turn a bad measurement into a confident chart. What AI changes is scale, distance, and form. It makes analysis nearly free, so far more numbers get produced and acted on, and it stacks layers, a model reading a summary of a dashboard built on a sensor, until the output sits several steps from the thing that was actually measured. Every extra layer is another place the source and its reliability can drop out. Some of them even hand off in prose rather than numbers, so each summary arrives already argued rather than measured, and working back down the chain means re-deriving arithmetic the sentences no longer contain. The problem is old. AI industrializes it.
I thought it’d be fun to analyze my watch data. The thing I actually came away with was about measurement, and it turned out to be more useful than anything I learned about my training, because I build software that reads a team’s messages and the question of what an agent is reasoning over is the same one at a larger size. My watch was covered by my sleeves, and I caught that because I knew what I’d been wearing. In a company, somebody knows about every sleeve, but it’s not always the person reading the summary.