Toto — the AI smart router that cuts your LLM spend

Toto

2 min read Original article ↗

Toto.

Route frontier. Build local.

Stop paying frontier prices for local work.

Toto routes each task to the cheapest capable model — frontier APIs when quality demands it, fine-tuned local models when it doesn't. We build the local models.

01 / 04 · San Francisco

Debug the risk-engine refresh path

deep reasoning · wide context

Summarize 200 counterparty emails

long context

Generate data-pipeline boilerplate

routine codegen

Classify every research note

domain-specific · high volume

Pre-screen comms for compliance

data never leaves perimeter

Toto Router

every decision logged, priced, auditable

Claude Opus$0.150/task

frontier reasoning · judgment-heavy work

Gemini 2.5 Pro$0.025/task

long-context summarization

DeepSeek V3$0.003/task

routine codegen at near-zero cost

Fine-tuned local≈$0.0004/task

your GPUs · your data never leaves

375× cheaper than frontier

Per run

Before $1.05

After $0.39

Saved $0.66 (63%)

Per year

Before $1.05M

After $390K

Saved $660K/yr (63%)

02 / 04

SSE · API · MCP · CLI Beta in production

03 / 04

04 / 04

FAQ · For teams cutting AI spend

Routing, local models, and your token bill.

What is an AI smart router?

An AI smart router sits between your tasks and the model market. It scores each incoming task and sends it to the cheapest model capable of doing the job — a frontier API for hard reasoning, a fine-tuned local model for routine patterns — instead of sending everything to one expensive model.

How much can routing cut our AI token spend?

Most teams send nearly every task to a frontier model by default and overpay for the routine ones. On our benchmark workload, Toto's routing cuts token cost about 63% with no loss in output quality — and the savings scale with task volume.

When does a task go to a local model instead of a frontier API?

When it's a pattern your workload repeats: classification, extraction, enrichment, templated drafting. Toto fine-tunes local models on those patterns. High-novelty or high-stakes tasks still escalate to frontier models.

Does Toto build the local models for us?

Yes. Toto builds, fine-tunes, evaluates, and maintains local models for your specific use cases from your task history. Your code and prompts never touch Toto's cloud — the models deploy where you control them.

How do we get started?

Toto is in private pilot. Drop your email on toto.tech or write to hello@toto.tech and we'll reach out.