GLM 5.3 Flash API pricing and speed | RunInfra

RunInfra

2 min read Original article ↗

GLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.11 per 1M input tokens and $0.45 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.

Pricing

USD, pay per token

per 1M input tokens
$0.11

per 1M cached input tokens
$0.03

per 1M output tokens
$0.45

Measured performance

Output speed

254.1output tokens per second, model only

Time to first token

703milliseconds to first reasoning token

Access

Confirm how your client reaches this model.

Provider
Z.ai

API compatibility
OpenAI-compatible chat completions
Anthropic compatibility
Anthropic-compatible Messages, POST /v1/messages
Accepted input
Text and images

Capacity

Check the limits your workload must fit.

Context window
1,048,576 tokens

Maximum request size
3.5 MB per request
Maximum generated output
1,048,576 tokens

Capabilities

See which request modes the API supports.

Tool calling
Supported

Precision
FP8, vendor-native release
View full spec

Gateway compatibility
OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways

Prefix caching
Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.

Cache retention
Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.

Upstream model
View model