GLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.11 per 1M input tokens and $0.45 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.
Pricing
USD, pay per token
- per 1M input tokens
- $0.11
- per 1M cached input tokens
- $0.03
- per 1M output tokens
- $0.45
Measured performance
Output speed
254.1output tokens per second, model only
Time to first token
703milliseconds to first reasoning token
Access
Confirm how your client reaches this model.
- Provider
- Z.ai
- API compatibility
- OpenAI-compatible chat completions
- Anthropic compatibility
- Anthropic-compatible Messages, POST /v1/messages
- Accepted input
- Text and images
Capacity
Check the limits your workload must fit.
- Context window
- 1,048,576 tokens
- Maximum request size
- 3.5 MB per request
- Maximum generated output
- 1,048,576 tokens
Capabilities
See which request modes the API supports.
- Tool calling
- Supported
- Precision
- FP8, vendor-native release
View full spec
- Gateway compatibility
- OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways
- Prefix caching
- Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.
- Cache retention
- Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.
- Upstream model
- View model