Lotus Root Is Not Fried Tofu

· Little Theta LLC ·

22 min read Original article ↗

field notes / on-device vision, part two

A user's photo logs came back wrong. The model turned out not to know what a fig looks like. Fixing that meant giving an on-device app a second set of eyes, on three platforms, without a server.

Qwen3-VL 4B · 2B llama.cpp b10107 · mtmd iOS in-process xcframework Android bundled server · Vulkan 14 plates · 4 models 1.6 GB phone · 3.3 GB desktop

Somebody took a picture of sliced lotus root on a grill pan, and my app told them it was eleven ounces of fried tofu. 842 calories. They logged it, because why wouldn't you, and then they sent me the screenshot.

I want to be honest about how that felt, because I think the feeling is the useful part. A week earlier I had written a whole field note about how photo food logging worked in Macro Trainer with no server, using the same 2-billion-parameter model that runs the coach. Eighteen of nineteen bench plates read sensibly. I was pleased with it. And here was a real person, with a real plate, getting confidently lied to by the thing I was pleased with.

This is the write-up of what happened next: the bench that proved the model was the problem and not the phone, the purpose-built vision model that replaced it, the number we had been throwing away that would have caught the tofu, and the two very different ways we got that model running on iPhones and Android phones. If you build anything that looks at pictures on a device, the middle third is for you.

Sliced lotus root, pale discs full of holes, laid out on a black grill pan with a bundle of shimeji mushrooms at the top.

The review screen on a Pixel: the same photo, then 'Lotus Root, Cooked, Boiled, Drained, Without Salt' at 1.25 × 10 slices and 'Mushrooms, White, Cooked' at 3.5 × 1 mushroom, 85 kcal total, with a Log 2 foods button.

The plate that started it, and the same plate read by the vision pack on the same phone two days later: lotus root and mushrooms, 85 calories, instead of fried tofu at 842. The user's photo, reproduced with their permission.

Where we left off

Quick recap for anyone who didn't read part one. Macro Trainer is a nutrition and training coach that lives entirely on your device. The coach chat runs on Gemma 4 E2B. Photo logging reused it: the phone builds send the photo through Google's LiteRT-LM runtime, the desktop builds through llama.cpp's llama-server with a separate projector file, the model answers with a tiny JSON list of foods and grams, a catalog matcher turns "bacon" into the right row of a million-row food database, and you check the result before anything is logged. The big finding in part one was that a small vision model has to be allowed to think first, or it glances at the plate and guesses.

So when the lotus root screenshot arrived, along with three more from the same person, my first thought was the obvious one. It's a phone. Phones have no JSON grammar. Maybe thinking got switched off somewhere on the Android path. Something in the plumbing.

The four logs that started this · one user, one grill pan, Android
What was on the plateWhat got logged
Sliced lotus rootTofu, Fried · 11 oz · 842 kcal
A tray of shimeji mushroomsChicken pieces · 428 kcal
Figs on red ricePotatoes + Frankfurter · 730 kcal
Grilled figs and peaches; tomatillos"Couldn't make out a meal"

It was not the plumbing

Here is the thing about having a bench harness already built: you get to find out you're wrong in an afternoon instead of a month. The user sent all fourteen of their photos, all grilled and roasted produce on the same pan. I ran them through the exact production pipeline on the Mac, where there is a real GPU, a JSON grammar, and thinking is definitely on. Best possible conditions for Gemma.

Four out of fourteen.

Lotus root came back as "white round crackers", 300 grams, 1,204 calories. Shimeji came back as "stir-fried vegetable strips", which the catalog then matched to macaroni. A platter of figs, berries, tomatillos, lotus root and mushrooms came back as pineapple, kiwi, apricot, dried fruit and nut clusters. None of those foods were in the picture. The model wasn't misreading the plate. It was making up a plate it had seen before.

Then I did the thing you're supposed to do, which is try the bigger model. Gemma 4 E4B is the tier we ship on 16 GB desktops, and it should know more. Two out of fourteen. The fig tray became roasted potatoes and sweet potato at 910 calories. The mushrooms became "cooked poultry", which the catalog turned into frankfurters.

The bigger Gemma wasn't smarter about food. It was a more confident liar.

That was the moment the diagnosis changed. Thinking helps a model organize what it saw. It cannot recover something it never learned to see. Gemma 4 is a very good general model with a vision encoder bolted on, and it simply does not know what a lotus root, a shimeji mushroom or a grilled fig look like. When it doesn't know, it does what language models do: it fills the gap with the most plausible-sounding thing. The failure mode of "confident and wrong" that I'd worried about in part one was the whole story, and no prompt was going to fix it.

A knowledge gap needs different knowledge. So the question became: is there a model that was actually trained to look at things, that fits in the runtime we already ship?

Borrowing better eyes

Qwen3-VL is a family of vision-language models, purpose-built for exactly this: look at an image, say what's in it. The 4B Instruct variant is about the same size on disk as the Gemma we were running, and llama.cpp had supported its projector since before the build we pin. I ran the same fourteen plates through it with nothing changed but the model files.

Twelve out of fourteen. In two to six seconds a plate, against Gemma's fourteen to twenty, because the Instruct build has no reasoning pass to pay for. It named the lotus root. It named the shimeji. It named figs, peaches, pears and lemons on a plate where Gemma had found "grilled peaches" and stopped. On the wide platter shot it listed ten items, all of them present, plus one invented "grilled chicken" that the calorie check I'll get to in a minute would later throw out.

Same 14 plates, same pipeline, four models · Apple Silicon, Metal
ModelFilesPer plateRead correctly
Gemma 4 E2B · thinking on (shipped)3.1 GB + 1.0 GB14–20 s4 of 14
Gemma 4 E4B · thinking onbigger30–67 s2 of 14
Qwen3-VL 4B Instruct Q4_K_M2.5 GB + 0.8 GB2–6 s12 of 14
Qwen3-VL 2B Instruct Q4_K_M1.1 GB + 0.4 GB1–4 s11 of 14
What each model said, before the catalog · a few of the plates
PlateGemma 4 E2BQwen3-VL 4B
Lotus root + shimeji"white round crackers" 300 g → 1,204 kcallotus root, mushrooms
Pear, blueberries, strawberries, mintblueberries, strawberries, "apple slices"all four
Kale"broccoli"broccoli, kale
Figs, peaches, pears, lemons"grilled peaches" onlyall four
Platter of six thingspineapple, kiwi, apricot, dried fruit, nut clustersfigs, blackberries, strawberries, tomatoes, lotus root, mushrooms
Broccolini, asparagus, celery, figsmixed veg, "grilled meat", potatoes, pistachiosbroccoli, asparagus, celery, figs
Tomatillos + purple sliceszucchini, red onioneggplant, green tomato

A loose scatter of grilled shimeji mushrooms, thin pale stems with small brown caps, on a black grill pan.

Halved figs, peach and pear halves and small green lemons arranged in rows on a grill pan.

A round wire tray holding bowls of grilled figs, blackberries, strawberries, tomatoes, lotus root and mushrooms.

Three of the fourteen. Left: Gemma said "stir-fried vegetable strips, button mushrooms, cooking oil" and the catalog logged macaroni; Qwen said grilled mushrooms. Middle: Gemma said "grilled peaches", 450 g of them; Qwen said fig, peach, pear, lemon. Right: Gemma said pineapple, kiwi, mixed berries, apricot, dried fruit mix and nut clusters, none of which are on the tray; Qwen said figs, blackberries, strawberries, tomatoes, lotus root, mushrooms. All three photos the user's, with permission.

I'll say the caveat out loud: fourteen plates from one grill pan is a narrow set. It is also a real set, from a real user, of exactly the kind of food that a general model has never been shown. The nineteen public plates from part one were bacon and eggs and cheeseburgers. Everybody's model knows a cheeseburger. The plates that matter are the ones that break it.

So the decision was made: on desktop, photos would be read by Qwen3-VL and the coach would stay on Gemma, which is still the better conversationalist. That combination has a name in the app now. It's called the vision pack.

What a vision pack actually is

Two files. The model weights as a quantized GGUF, and the multimodal projector, the mmproj, which is the part that turns image patches into something the language model can read. On desktop that's the 4B at about 3.3 GB together. On phones it's the 2B at about 1.6 GB. Both live in the same folder as the coach model, and the app's model-directory cleanup was taught which filenames it owns, because I learned the hard way on day one that a cleanup routine which deletes "any GGUF the spec doesn't recognize" will happily delete a projector while you're in the middle of benchmarking it.

The interesting engineering is not the files. It's that desktop runs one resident llama-server, one model at a time, and now two different models want it.

photo camera · library · file encode ≤1024 px JPEG ≤320 image tokens the vision pack Qwen3-VL + mmproj desktop: llama-server, borrowed iPhone: in-process llama.cpp Android: bundled server, Vulkan foods JSON name · grams · kcal protein · carbs · fat catalog match FTS5, generic-first pool whole-word ranker + calorie density check best household serving review screen edit amount · swap match · untick log as ordinary diary rows, photo attached

Borrowing the server

In part one the projector was part of the server's identity: change it and the bridge restarts the process. That still holds. What's new is that a photo now restarts the server under a completely different model, and the coach, which was mid-conversation in that process with a KV cache full of transcript, has no idea it happened. The Swift side on macOS can restart the process without the Dart engine ever seeing a signal.

The fix is a counter. There's a tiny static class I named ResidentServerTenancy, which is a grand name for an integer. The recognizer bumps it after every desktop pass, success or failure, in a finally. The coach engine snapshots it when it starts a turn and compares it when the turn ends. If the number moved, the engine invalidates its session and throws the same "session unavailable" error a crashed server would, which lands in the recovery path that already existed: re-send the whole transcript and carry on. Nothing is hidden, nothing is duplicated, and the photo pays a few seconds of model load that the user can see.

resident_server_tenancy.dart · the whole mechanism

/// Bumped by anything that restarts the shared llama-server under a
/// different model. The coach compares generations across a turn and
/// treats a change as "my session is gone", which it already knew how
/// to survive.
class ResidentServerTenancy {
  static int generation = 0;
  static void serverBorrowed() => generation++;
}

// meal_photo_recognizer.dart, desktop path
try {
  return await _recognize(bytes, serverPath, pack.modelPath, pack.projectorPath);
} finally {
  ResidentServerTenancy.serverBorrowed();   // even on failure: the server is Qwen's now
}

One side effect I'm fond of: the coach no longer loads a projector at all. In part one every caller had to pass the Gemma projector on every launch so the server wouldn't restart. Now the coach launches text-only, the photo pass is the only thing that ever loads a projector, and the desktop coach uses about a gigabyte less memory than it did last week. Replacing a feature is sometimes the cheapest way to simplify it.

The number we were throwing away

The tofu wasn't only the model's fault, and this is the part I'd underline if you're building anything like it. Go back to that first log. The phone model said "tofu", not "fried tofu". The catalog took "tofu" and upgraded it to "Tofu, Fried" on the way to the diary, because the ranker had a small bonus for cooked entries and "fried" is a cooking word. The user's plate had no oil on it. The model had actually said so, in a way.

Every answer from the model carries a calorie estimate for the portion. We had been parsing it, keeping it in case the catalog had no match, and otherwise ignoring it. Meanwhile the catalog hit we chose had its own calories per 100 grams, right there in the row. Two numbers, describing the same food, and we never compared them.

Now we do. Each candidate hit's density is checked against the model's own kcal-per-100-grams. Within a factor of two is free, because portion guesses are rough. Beyond that the penalty grows with the log of the ratio, about four points at four times off, capped at six. That's enough to sink a wrong food without overriding a right one. And both Qwen's and Gemma's density guesses turned out to be reliable to within that factor of two for whatever they named, which is the only reason this works at all.

meal_photo_matcher.dart · the density check

/// Free within 2× of the model's own estimate, then 6 points per
/// factor of e beyond that, capped at 6. Wrong foods are usually
/// wrong by a lot: a pickle jar at 91 kcal against zucchini at 17.
double calorieDensityPenalty(double hitKcalPer100, double modelKcalPer100) {
  final ratio = max(hitKcalPer100, modelKcalPer100) /
                max(min(hitKcalPer100, modelKcalPer100), 1);
  if (ratio <= 2) return 0;
  return min(6, 6 * log(ratio / 2));
}
Matches the density check settled · all from the same fourteen plates
Model saidRanker pickedNow
TofuTofu, FriedTofu, Firm
Grilled zucchiniZucchini Grilled (a jar, 91 kcal/100 g)Squash, Zucchini, Cooked
Fresh mintFresh Mints (the Tic Tac)Spearmint, Fresh
Grilled tomatoesa 5 kcal Waffle House sideTomatoes, Red, Cooked

Three smaller rules rode along. "Fried", "breaded" and "battered" now cost points when the photo didn't ask for them instead of earning them; pan-fried stays ordinary cooking, because bacon. A query that itself says "grilled" no longer carries the raw-produce preference from part one, which had been sending grilled zucchini to raw zucchini. And mint got an alias to its species, because the catalog thinks in Latin-adjacent USDA names and the model thinks in English. The mint one was thirty seconds of work and I'd been shipping a breath mint for a garnish.

We had the one number that would have caught every bad match, and we were parsing it just to discard it.

Phones are where it gets ugly

Desktop took a day. Phones were the real project, because the phone builds don't run llama.cpp at all. They run LiteRT-LM, and LiteRT-LM runs one file format, .litertlm, and there is no Qwen3-VL in that format. If phones were going to get the better eyes, they were going to get a second runtime to open them with. Two platforms, two answers, neither of them the same.

iPhone: put llama.cpp inside the app

llama.cpp publishes an xcframework with every release, Metal on, multimodal library included. I wrapped it as a local CocoaPod so it links and embeds without hand-editing the Xcode project, wrote a Swift bridge over the C API, and ran the whole thing in-process. Load the model, load the projector, build the chat prompt with the media marker where the image goes, evaluate the image chunks through the helper, then greedy-decode tokens until a small state machine sees the JSON close. No server, no port, no process to supervise.

On an iPhone SE, the four-gigabyte floor device, a plate reads in seven to twelve seconds after a five-second load. Peak memory is about 920 MB. That's better than the twenty-five seconds Gemma took on the same phone in part one, with a model that actually knows what it's looking at. I was not expecting the small phone to be the good news.

Three things bit, in order of how long they cost me. Image tokens: a wide shot of a tray came out to 609 image tokens and the OS killed the app on the spot, so the encoder is capped at 320 and a wide plate gets read a little coarser rather than not at all. The simulator: Metal in the iOS simulator cannot allocate the projector's buffers, so under targetEnvironment(simulator) it runs on the CPU and you stop trusting simulator timings. And the phone itself: a locked iPhone kills GPU apps and refuses launches, and a dropped debugger console wedged the developer disk image so thoroughly that the only thing that cleared it was toggling Developer Mode off and on and rebooting. I lost an evening to that one. Write it down somewhere you'll find it.

LlamaVisionBridge.swift · the decode loop, minus the ceremony

let stop = JsonStop()                        // depth counter + repeat guard
var out = ""
while out.utf8.count < maxBytes {
  let logits = llama_get_logits_ith(ctx, -1)
  let tok = argmax(logits, nVocab)          // greedy; temperature 0 on every platform
  if llama_vocab_is_eog(vocab, tok) { break }
  let piece = tokenToPiece(tok)
  out += piece
  if stop.feed(piece) { break }            // outermost brace closed, or the model is looping
  batch = llama_batch_get_one(&tok, 1)
  guard llama_decode(ctx, batch) == 0 else { break }
}

Android: ship the server as a fake library

Android has no in-process bridge, and I didn't write one, because the app already had something better: the Dart-side llama-server bridge that Linux and Windows use, with its loopback HTTP client, its JSON grammar, and a parent-death watchdog that kills the child if the app dies. All Android needed was the executable.

So the Android build cross-compiles llama-server with the NDK and drops it into jniLibs/arm64-v8a/libllama_server.so. It is not a shared library. It's a static executable wearing a .so extension so that Gradle packages it and, with legacy packaging turned on, Android extracts it to the native library directory at install time, where the app is allowed to execute it. The Dart bridge spawns it through /system/bin/sh under the same watchdog it uses on Linux. Two small tolls: the server's tool runner needs posix_spawn, which Android's libc only grew at API 28, so the pack is gated to Android 9 and up; and Play now wants native libraries 16 KB page-aligned, which the NDK 28 toolchain does by default and which I checked with llvm-readelf before believing.

Then I benchmarked it on a Pixel 8 Pro, and this is where I'd like you to picture me making a face.

Qwen3-VL 2B on a Pixel 8 Pro · Tensor G3, Mali G715 · same three plates
BuildTypical plateWhat's going on
CPU, cold phone29 sFine. Once.
CPU, third plate45–85 sThe SoC throttles and decode collapses to 1.5–2 tokens a second. The phone is warm to the touch.
Vulkan, weights on GPU35–40 sDecode steady at 11–12 tok/s across every run, no throttling. Prefill is now the wall: 25–32 s for the image tokens on the Mali's Vulkan kernels.
Vulkan, Q8_0 weights34–41 sPrefill identical. Decode halves to 7 tok/s, because it's memory-bound and the weights are twice the bytes. A wash.
Vulkan, llama.cpp 760 commits newer36–63 sPrefill identical again. The ceiling is the GPU's kernels for this shader path, not the version.
Gemma via LiteRT-LM (before)~22 sFaster, and wrong. Mostly wrong.

I tried everything cheap. Batch sizes: no change. Halving the image tokens: halved prefill and hallucinated chicken onto a plate of berries, so no. Q8 weights on the theory that dequantizing Q4 was the cost: prefill didn't move an inch, which proved it wasn't. A newer llama.cpp: same wall. The Mali's Vulkan matrix multiply is what it is, and a 2B model's image prefill on it is thirty seconds. The CPU can do the prefill faster when it's cold, at 37 tokens a second, and then it cooks itself.

So the Android number is about forty seconds a plate. That's a product decision, not an engineering one, and I made it the way I'd want it made for me: ship it, say so, and don't hide it. The offer card on Android literally says "A plate takes about 40 seconds to read." The gate requires a Vulkan 1.3 GPU on top of six gigabytes of memory, which in practice means Android 13 or newer, because a phone without Vulkan would fall back to the CPU path and the CPU path is the one that throttles. Forty seconds and right beats twenty-two seconds and lotus-root-is-tofu. I've seen the screenshot.

2–6 sper plate · Mac · Metal · 4B

7–12 sper plate · iPhone SE · 2B

35–40 sper plate · Pixel 8 Pro · Vulkan · 2B

920 MBpeak on the SE, in-process

12 / 14plates the pack reads · vs 4

0 bytesleave the device

Why the newer phones don't get the "better" model

A reasonable question came up while all this was going on: newer phones have more memory, so shouldn't they get the Q8 weights, and older ones the Q4? The bench answered it, and the answer generalizes, so I'll spell it out.

Decoding is bound by memory bandwidth. Every token, the whole weight file streams through the processor once. Q8 is twice the bytes of Q4, so it decodes at half the speed on any phone, however new. A faster phone decodes both faster, but the ratio stays. And what Q8 buys you on a 2B model is small: tidier JSON on one plate. The lever that actually moves accuracy is the 4B model, which is the desktop pack, and which would fit an 8 GB iPhone comfortably. That's a tier worth building when there's a phone on the bench to prove it on. Quantization isn't. One pack per platform, Q4, until the numbers say otherwise.

What shipped

Desktop got the 4B pack first: the photo screen offers a one-time 3.3 GB download, the coach stays on Gemma, and photos are read in a few seconds on Apple Silicon and somewhat longer on Windows and Linux where it's CPU. Then iPhone got the 2B pack as an option, offered on the photo screen to anything with three gigabytes of memory or more, with a Remove button in the coach settings, and photos still work without it. Android got the same offer one release later, behind the Vulkan gate, with the forty-second warning in the card.

The photo logging screen on a Pixel 8 Pro. Above the camera button, a card titled 'Sharper photo recognition' explains the 1.6 GB download, says photos never leave the phone and that a plate takes about 40 seconds to read, with 'Not now' and 'Download 1.6 GB' buttons.

The offer on the Pixel 8 Pro, from the build that shipped. The forty seconds is in the copy on purpose. The download is resumable, the pack lives beside the coach model, and removing it from settings puts photos back on the coach's own eyes.

The cut itself found two more things that had nothing to do with vision and everything to do with shipping, and I'll mention them because they're the kind of thing a solo developer hits at eleven at night. The iOS xcframework and the Android server binary are both ignored by git, on purpose, and both were only ever staged by hand on my Mac. The first release build after the spike would have failed to link on iOS. Now the TestFlight job downloads the framework from the pinned llama.cpp release, and the Play job downloads the server binary from a release on our own repo, checksum-verified, the same way it already fetched the food database. If your CI never built the thing you benchmarked, your CI has not built the thing you benchmarked.

What I'd tell you before you build one

  1. Bench the failures, not the successes. Nineteen public plates said the feature worked. Fourteen plates from one user's grill pan said the model didn't know what a fig was. The plates that break it are the only ones that teach you anything.
  2. Separate "can't reason" from "can't see". Thinking fixes the first. Nothing but a different model fixes the second, and the bigger version of the same model just lies with more confidence.
  3. Use the number the model already gave you. The calorie estimate was in every answer and we discarded it. Two independent estimates of the same thing are a free sanity check. Look for the field you're parsing and ignoring.
  4. A shared resident model needs a tenancy signal. A counter and a "session gone" error that your recovery path already understands. Don't try to hide the restart; make it loud and cheap.
  5. On phones, the runtime is the decision. LiteRT-LM had the model we started with and not the one we needed. In-process llama.cpp on iOS and a bundled server on Android were the fastest honest paths, and they're different because the platforms are.
  6. Cap the image tokens before the OS caps your process. 609 image tokens was a SIGKILL on a 4 GB phone. 320 is a slightly coarser read. Take the coarser read.
  7. Time prefill and decode separately, then believe the timing. Q8 and a newer build both changed nothing about the Mali's prefill, and the bench said so in an hour each. That's what stopped me from spending a week on it.
  8. Ship the slow-and-right one, and put the number in the UI. If it takes forty seconds, the card should say forty seconds. People forgive slow. They don't forgive tofu.
  9. If it's gitignored and CI needs it, CI has to fetch it. A release pipeline that has never built your native artifacts will tell you so at the worst possible moment.

Part one said the model was never the interesting part. Part one was wrong about that for exactly one week.