Where My LLM API Bill Actually Went — and How I Cut It

· Medium ·

9 min read Original article ↗

A practical breakdown of the architecture changes that reduced my coding-agent costs.

Colin Peng

Press enter or click to view image in full size

My LLM bill was not high simply because the models were expensive. It was high because I was:

  • Sending the same 1,800-token system prompt on every call
  • Retrying failures immediately, without backoff
  • Defaulting every workflow step to a frontier model, whether it needed one or not

In my workload, I estimated that headline price per token explained roughly 20% of the problem. My own architecture explained the other 80%.

That split is specific to my system, not a universal benchmark. But the underlying lesson is broadly useful: the listed token price tells you the rate. It does not tell you how efficiently your application uses that rate.

Before changing providers or negotiating prices, I instrumented my spending. That was the step that made every later decision evidence-based rather than intuitive.

You cannot optimize what you do not log.

Step 1: Attribute Every Token to a Workflow Step

The first requirement for LLM cost optimization is step-level attribution.

For every API call, I recorded:

  • The workflow step
  • The model
  • Input tokens
  • Output tokens

I then aggregated the data nightly by step and model.

When prices are quoted per million tokens, the cost of a workflow step can be calculated as:

step cost =
number of calls × [
(average input tokens ÷ 1,000,000 × input price per 1M tokens)
+
(average output tokens ÷ 1,000,000 × output price per 1M tokens)
]

After one week, I had a representative breakdown of my coding-agent workload.

The most important finding was unexpected: the planning step accounted for only 8% of calls but 41% of total cost.

Planning was expensive because it combined the longest contexts with the most expensive model. Before measuring, I had assumed code generation was the costly part. The data showed that my intuition was wrong.

The relevant metric was not request volume alone. It was total spend per workflow step.

A Minimal Implementation

The implementation was simple because the API response already included token-usage data:

import logging
from openai import OpenAI

client = OpenAI(
api_key="KEY",
base_url="https://api.cometapi.com/v1",
)
def tracked_call(step_name, model, messages):
response = client.chat.completions.create(
model=model,
messages=messages,
temperature=0,
)
usage = response.usage
logging.info(
"llm_call step=%s model=%s prompt=%d completion=%d",
step_name,
model,
usage.prompt_tokens,
usage.completion_tokens,
)
return response

Once these logs exist, the cost report is essentially a GROUP BY operation.

The important part is not building an elaborate observability platform. It is ensuring that every token can be attributed to both a workflow step and a model.

Step 2: Take the Gateway Saving, but Treat It as a Baseline

I route requests through CometAPI. In my usage, the resulting cost was roughly 20% below what I would have paid through direct provider APIs for the same model mix.

This was a low-effort saving. I could retain the same client pattern and model choices while reducing the invoice across the models I was already using.

I took the saving immediately because it required almost no architectural work.

But the gateway discount was not the largest lever. The architecture changes that followed saved more, although they required actual engineering.

This distinction matters because implementation order and financial impact are not always the same:

  • Take low-effort pricing savings immediately.
  • Spend engineering time on model routing, prompt overhead, and retry behavior.

There is little reason to postpone an immediate saving while building a more sophisticated optimization layer. At the same time, a lower API rate cannot compensate for an inefficient request architecture.

CometAPI explains its billing model in its pricing documentation.

Step 3: Route Each Workflow Step to the Right Model Tier

Model tiering means assigning different models to different workflow steps according to the quality each step actually requires.

My planning step genuinely needed a frontier model. It was low-volume but high-leverage: a poor plan could degrade every downstream action.

Bulk edits and lint fixes did not have the same requirements. They were high-volume, more constrained, and easier to retry or verify.

I changed the routing logic accordingly:

def choose_model(step):
if step in ("plan", "hard_refactor"):
return "claude-sonnet-4-6" # High-value reasoning steps

if step in ("bulk_edit", "lint_fix"):
return "gemini-2.5-flash" # Lower-cost, fast, retry failures
return "gpt-5-mini" # Middle-tier default

This moved 46% of my calls to a cheaper, faster model while reserving the expensive model for planning and difficult generation tasks.

Model tiering produced the largest single line-item reduction in my bill. It worked because it targeted the highest-volume workflow steps with cheaper models rather than trying to save small amounts on already low-volume steps.

The principle is straightforward: use the most capable model where errors have the highest downstream cost, not everywhere by default.

Step 4: Remove Repeated Prompt Overhead

My 1,800-token system prompt was being sent with every request.

Much of it was defensive boilerplate that had accumulated over months. Individual instructions had seemed harmless when added, but together they created a substantial fixed cost on every call.

I removed redundant instructions, kept only the essential constraints, and moved the stable portion into a cached prefix where the provider supported prompt caching.

The scale of repeated prompt overhead is easy to underestimate. At one million calls, an unchanged 1,800-token system prompt alone would represent:

1,000,000 calls × 1,800 tokens
= 1.8 billion input tokens

That is before counting any task-specific context.

Prompt tokens are still tokens. In my system, the oversized prompt increased both API cost and consumption of the tokens-per-minute allowance on every request.

Prompt cleanup therefore affected two dimensions:

  • It reduced recurring input-token cost.
  • It increased the amount of useful work that could fit within the same rate limits.

Caching helped with the stable portion, but trimming came first. Caching unnecessary text would have preserved the underlying prompt-design problem.

Step 5: Replace Blind Retries with Backoff and Jitter

Retries accounted for 12% of my calls.

Many followed an HTTP 429 rate-limit response. The system would retry immediately under nearly identical conditions, receive another 429, and sometimes appear to prolong the rate-limited period.

I replaced immediate retries with exponential backoff and jitter.

Exponential backoff progressively increases the delay between attempts.

Jitter adds a small amount of randomness so concurrent requests do not all retry at exactly the same time.

This reduced retry volume by more than half.

Not every rejected request produces the same billing outcome. However, blind retries still increase request volume and can trigger repeated billable work when requests reach model execution.

In my workload, retry control was therefore both a reliability improvement and a cost-control measure.

The broader rule is simple: a retry should respond to the cause of failure. Repeating the same request immediately is not a recovery strategy.

Which Changes Reduced Cost Most

The four changes affected different parts of the cost equation.

Model tiering produced the largest single line-item reduction. It moved 46% of calls to a cheaper model while preserving the expensive model for high-leverage work. Implementing the change took about one day.

Gateway routing provided the easiest immediate saving. In my usage, it reduced the cost of calling the same model mix by roughly 20% and required almost no engineering work.

Prompt trimming and caching removed recurring input overhead. The original 1,800-token system prompt had been attached to every request, regardless of how simple the task was.

Backoff and jitter reduced unnecessary request volume. Retries had accounted for 12% of all calls, and the new retry logic cut that volume by more than half.

The combined effect exceeded what any single change could have achieved.

These savings should not be added as though they were independent discounts. Each intervention changes a different part of the system:

  • Gateway pricing changes the rate paid for tokens.
  • Model tiering changes the blended rate across workflow steps.
  • Prompt trimming changes the number of input tokens.
  • Retry control changes the number of requests.

Each change also alters the cost base on which the next change operates.

How to Calculate Monthly LLM API Cost

Suppose a system processes one million requests per month, averaging 2,000 input tokens and 500 output tokens per request.

Its monthly token volume is:

Input:
1,000,000 × 2,000
= 2 billion tokens

Output:
1,000,000 × 500
= 500 million tokens

Using prices quoted per million tokens, the baseline monthly cost is:

monthly cost =
2,000 × input price per 1M tokens
+
500 × output price per 1M tokens

In my case, routing through the gateway reduced that baseline by roughly 20% before any architectural changes.

Model tiering then changed the blended token rate because 46% of requests moved to a model that cost a fraction as much per token. Prompt trimming reduced the input-token volume on every affected request. Retry control reduced unnecessary request volume.

The correct calculation therefore depends on four variables:

  1. Request volume
  2. Input and output token mix
  3. Model allocation by workflow step
  4. Current prices for each model

That is why I do not quote a universal “you will save X%” figure. The only useful estimate is one calculated from your own workload.

Three Failure Modes of LLM Cost Optimization

Optimizing Low-Volume Steps Because They Are Easy

It is easy to spend time shaving pennies from a low-volume step while a high-volume step continues consuming most of the budget.

Optimization targets should be sorted by total spend, not by how simple they are to change. This is why step-level cost attribution must come first.

Cutting Cost Without Testing Quality

A workflow step moves to a cheaper model, the invoice falls, and the change appears successful. Three weeks later, someone notices that the output has degraded.

Every model substitution should be gated by an evaluation set: a representative collection of tasks used to detect quality regressions.

A cost reduction that fails the quality threshold is not a successful optimization.

Building a Cache Before Measuring Reuse

Caching sounds like an obvious cost solution, but it only helps where content actually repeats.

An elaborate caching layer built before measuring prompt reuse can consume engineering time without addressing the dominant cost driver.

Measure repetition first, then cache the prefixes or responses that genuinely recur.

Cost optimization performed without measurement can improve the invoice while making quality, latency, or reliability worse.

Keep Cost Instrumentation Running

Instrumentation should remain active after the first optimization pass.

LLM costs drift over time:

  • A system prompt gradually grows.
  • A new workflow step defaults to the expensive model.
  • A retry bug returns.
  • Request distribution shifts toward a previously minor step.

The dashboard that identifies the original problem is also what prevents it from silently returning.

I review my step-level cost table once a month. The review takes about ten minutes and has already caught two regressions.

LLM Cost Is an Architecture Problem

Price per token is only the rate. Your architecture determines how many tokens you buy, which models process them, and how often work is repeated.

In my coding-agent workload, the most effective sequence was:

  1. Instrument every call by workflow step and model.
  2. Take the low-effort gateway saving.
  3. Route each step to the cheapest model that meets its quality requirements.
  4. Remove repeated prompt overhead and cache stable prefixes where supported.
  5. Replace blind retries with backoff and jitter.
  6. Keep monitoring permanently so costs do not drift back.

Model tiering delivered the largest single reduction. Gateway routing produced an immediate saving of roughly 20% in my usage. Prompt cleanup removed fixed overhead from every request, and better retry logic cut retry volume by more than half.

Chasing the lowest advertised token price is far less useful than understanding your own token flow.

LLM cost is an architecture problem wearing a pricing costume — and measurement is what makes that architecture visible.