Are API providers secretly serving us lobotomized models or just burning VC cash at a suicidal rate? Because the inference math on @TheMoonshotAI Kimi K2.5 makes absolutely zero sense.
I just benchmarked the actual nvidia/Kimi-K2.5-NVFP4 weights on an 8x RTX PRO 6000 Blackwell cluster. Keeping this rig alive on @runpod (which is fairly cheap compared to hyperscalers) costs $13.52/hour (nearly $10k/month).
Looking at my raw baseline numbers for single requests:
Short text (40 in, 64 out):
4.9s end-to-end -> $0.0184 raw compute cost
Long multimodal (3.8k tokens + image):
3.0s end-to-end -> $0.0112 raw compute cost
Meanwhile, you can hit the official API for dirt cheap—just $0.60 per million input tokens
1. How the hell are they offering it so cheaply? Even if we assume god-tier continuous batching, extreme multi-tenant saturation, and enterprise hardware discounts, the hardware overhead vs. API revenue gap is astronomical.
Something isn't adding up. Either we are being served aggressively pruned, "diet" weights behind closed doors disguised as the full model, or there's some dark magic happening in their serving stack that I don't know about.
I've dropped my full vLLM configurations and throughput results below. Infra folks, please tell me what I'm missing. Am I completely botching the runtime setup, or are we simply not getting the real deal via the API? 👇
Kimi K2.5 NVFP4 on 8x RTX PRO 6000 on RunPod
This report compares two RunPod deployments of nvidia/Kimi-K2.5-NVFP4 on identical hardware:
- baseline pod keleoo019f2saw
- latency_tuned pod tlu3x0h0vvno4r
Both used:
- 8x RTX PRO 6000 Blackwell Server Edition
- RunPod region EUR-IS-2
- vllm/vllm-openai:latest
- tensor-parallel-size 8
- mm-encoder-tp-mode data
What changed between the two deployments
Baseline
The baseline deployment used the stable serving configuration I had already verified for streaming image input.
Key characteristics:
- standard vLLM scheduling behavior
- max-num-seqs 4
- default multimodal processor cache behavior
- default custom all-reduce behavior
Latency-tuned
The latency-tuned deployment kept the same model and hardware, but changed the runtime to favor single-request interactivity:
- --performance-mode interactivity
- --max-num-seqs 1
- --disable-custom-all-reduce
- --mm-processor-cache-gb 0
The goal was to see whether these settings improved streaming latency enough to justify the lower batch/concurrency posture.
Test method
I ran four single-request benchmark cases on each deployment:
- short text-only prompt
- short text + image prompt
- long-context text-only prompt
- long-context text + image prompt
Notes:
- For latency, I used streaming requests and recorded both time to first content token and total stream time.
- For exact token counts, I repeated each case once without streaming and used the returned usage object.
- The "long context" case is heavier than the user's example. Instead of stopping at about 1000 prompt tokens, I pushed a larger prompt and measured the actual tokenizer result. The text-only long prompt measured 3374 prompt tokens, and the text + image long prompt measured 3803 prompt tokens.
- The image case used the same public car image URL in both runs.
- These are single-request measurements, not saturation or multi-tenant throughput tests.
Cost
Observed runtime price from the live RunPod pod metadata for both deployments:
- 13.52 USD/hour
That implies:
- 324.48 USD/day if left running continuously
- 9869.60 USD/month at 730 hours/month
Per-request compute cost in the tables below is calculated from the observed runtime price and the measured total stream time.
Results
Takeaway:
- The tuned deployment was slightly faster across the board for the short multimodal case.
- The gains were real but modest: about 50 ms faster TTFC and about 216 ms faster total time.
Side-by-side summary
Where baseline is better
- much safer first-request behavior
- best short text-only experience
- best long-context text-only throughput
- higher prefill throughput on both long-context cases
Where latency_tuned is better
- best short multimodal request latency
- slightly faster total time on the long multimodal case
- better decode speed on the short image and long image cases
Sources
- RunPod Pod pricing: https://docs.runpod.io/pods/pricing
- RunPod getting started guide, stopped-pod cleanup note: https://docs.runpod.io/get-started
- NVIDIA model card for nvidia/Kimi-K2.5-NVFP4: https://huggingface.co/nvidia/Kimi-K2.5-NVFP4
- vLLM Kimi K2.5 recipe: https://docs.vllm.ai/projects/recipes/en/latest/moonshotai/Kimi-K2.5.html

