we spent september rewriting the kernels behind GLM 5.3 Flash. major release is live on RunInfra today 670 tok/s on Vercel AI Gateway $0.11 per 1M input, $0.45 per 1M output, $0.03 per 1M cached. 1M token context. FP8, vendor-native release it now runs on AMD. same model, same API, more capacity behind it 99.7% cache hit rate over the last 24 hours. every hit and miss shows up in your dashboard, so you can track the cache logs yourself OpenAI-compatible chat completions and Anthropic-compatible /v1/messages. text and image input, tool calling, JSON mode, streaming zero data retention. never used for training runinfra.ai/inference-api/…
