Gateway | VLM Run

2 min read Original article ↗

One unified API, Any Visual Model.

Document OCR, captioning, and multi-modal chat: every visual capability behind one MCP server.

22 visual models. One endpoint.

Pricing calculator

Dirt-cheap Document OCR.

Pick a document type, page volume, and OCR model. See cost savings versus closed vision APIs.

Gateway OCR tokens (in / out)

120M - 250M / 100M

Frontier billed tokens (in / out)

250M / 200M

2.5K image tokens per page; output includes reasoning at 2× OCR text.

Total costwhat the gateway bills for this workload

$23 - $118

Documents / month

100Kpages

Cost comparison (log scale)

Cheapest - most expensive

VLM Run Gateway

Document OCR VLMs

$23 - $118

OCR APIs

Textract, Azure Doc AI

$150 - $1K

Document AI APIs

Reducto, LlamaParse, Extend

$1K - $6K

Frontier VLMs

Gemini, Claude, GPT

$938 - $7.3K

Estimates vs typical OCR, Document AI, and frontier VLM pricing.Source: llm-prices.com

A vision-only gateway, built for builders.

LLM routers and gateways route to 100s of LLMs, yet only a handful of VLMs, and often no OCR or classical CV models. Visual AI deserves its own stack.

  • Multimodal, Multitask

    One catalog spanning multi-modal inputs and multi-task outputs: OCR, detection, segmentation, pose, keypoints, and more.

  • Chat Completions Native

    OCR, captioning, and multimodal chat run through the same OpenAI-compatible chat completions API you already use.

  • Orchestration Built-In

    Send a 500-page PDF or a 2-hour video in a single call. The gateway chunks, batches, and reassembles for you. No pipelines to build.

  • Structured Outputs, Out of the Box

    JSON-schema enforcement on every call. CV wrappers emit fixed schemas; VLMs honor response_format.

  • Open Models, No Lock-In

    Every model is open-weight, served through the OpenAI-compatible API you already use. Swap models or providers freely, your client code never changes.

  • Agent-native vision, via MCP

    Give any MCP client instant access to the full visual model catalog. Agents see, read, and reason over images out of the box.

Built for production visual AI.

Optimized for production workloads and cost-efficiency. Every model deployed gets its own performance tune-up.

  • 22

  • <100ms

  • 99.9%

  • SOC 2

Use cases

What you can build with the Gateway.

Compose vision models like building blocks. One SDK. One key. One bill.