Renting inference hardware in 2026 is getting weird in an interesting way.
There are more serious GPUs to choose from than there used to be. Some are older but still cheap and familiar. Some are newer and clearly faster in the right conditions. Some look amazing on a spec sheet and then behave very differently once a real serving stack, a real checkpoint, and a real cloud image get involved. The result is that people are no longer just asking, “Which GPU is fastest?” They are asking, “Which setup gives me the best tokens, latency, and cost profile for the model I actually want to serve?”
That is why token economics matter more now. Raw peak performance is nice. What actually hits the budget is the messy combination of TTFT, sustained decode, engine stability, and how many usable tokens fall out per infrastructure dollar. And because there are so many newer GPUs in the mix, and because performance changes from setup to setup, these comparisons are getting more valuable, not less.
This report is meant to help with exactly that. It covers a live RunPod benchmark run for the exact hardware shapes requested:
- 8x A100-SXM4-80GB
- 8x RTX PRO 6000 Blackwell Server Edition
Runtimes tested:
- vLLM
- SGLang
Model under test:
- moonshotai/Kimi-K2.5
The first surprise was that the older, cheaper A100 box felt faster at the beginning of a request.
The second surprise was that the newer, pricier RTX PRO 6000 box made up the difference and then some once decode speed and token economics entered the room.
The third surprise was that SGLang on both systems did the theatrical part perfectly, loading the whole model into VRAM, and then refused to do the practical part, actually serving traffic.
So the hardware stories diverged immediately:
- A100 was the faster first-token machine
- RTX PRO 6000 was the faster sustained-throughput machine
- SGLang on both boxes managed the dramatic part, loading the whole model into VRAM, and then forgot the crucial part, opening port 8000
- 8x A100 + vLLM served successfully
- 8x RTX PRO 6000 + vLLM served successfully
- 8x A100 + SGLang loaded the full checkpoint and sat around 76.7 GiB per GPU
- 8x RTX PRO 6000 + SGLang loaded the full checkpoint and sat around 76.1 GiB per GPU
So this was not a simple VRAM shortfall. The native INT4 checkpoint fit on both eight-way boxes. The real split was how the engine and the full deployment setup behaved after load.
Cost First
Before getting lost in token speed charts, here is the plain infrastructure price, because this is where most real-world decisions start:
A100 was cheaper by $1.60/hr. That sounds small, but the interesting twist is that the cheaper box won on TTFT while the more expensive one won on sustained speed and tokens per dollar.
That is not the usual shape of these stories, and it is exactly why setup-to-setup benchmarking is worth doing instead of assuming the newer card automatically wins every category.
This is where A100 looked better:
- short prompt TTFT: A100: 1.549 s PRO6000: 3.197 s
- long-context TTFT: A100: 0.364 s PRO6000: 0.599 s
So if your definition of “feels fast” is “the model starts talking quickly,” the A100 lane won both prompt shapes.
This is where the story flips.
- short prompt total latency: A100: 10.537 s PRO6000: 8.080 s
- long-context total latency: A100: 5.497 s PRO6000: 3.670 s
So the A100 got the first token out sooner, but the PRO6000 finished the job sooner.
The simplest analogy is:
A100 was the sprinter off the blocks. PRO6000 was the stronger cyclist over the full course.
The sustained decode numbers were not close.
Short prompt:
- 8x A100 + vLLM prefill: 14.20 tok/s decode: 6.12 tok/s end-to-end: 7.31 tok/s
- 8x RTX PRO 6000 + vLLM prefill: 6.88 tok/s decode: 13.11 tok/s end-to-end: 10.64 tok/s
Long-context prompt:
- 8x A100 + vLLM prefill: 2612.64 tok/s decode: 10.91 tok/s end-to-end: 183.19 tok/s
- 8x RTX PRO 6000 + vLLM prefill: 1587.65 tok/s decode: 22.14 tok/s end-to-end: 277.66 tok/s
Takeaway:
- A100 had the stronger prefill path
- PRO6000 had roughly 2x the decode speed
That decode lead is exactly why the PRO6000 overtook A100 on total latency despite losing on TTFT.
Generated Tokens per Dollar
This is the part I expected to be tighter, because the A100 box was cheaper and the hourly delta was not huge.
It was not tighter.
Generated completion tokens per USD:
- short prompt: A100: 1,576.42 tok/USD PRO6000: 2,109.09 tok/USD
- long-context prompt: A100: 3,076.72 tok/USD PRO6000: 4,933.65 tok/USD
That means:
- PRO6000 was about 1.34x better on short-prompt token economics
- PRO6000 was about 1.60x better on long-context token economics
So even though the PRO6000 box cost more per hour, it extracted enough extra decode throughput to win the value race too. That is the kind of result that matters in practice, because tokens per dollar are quietly becoming one of the most useful reality checks in inference benchmarking.
The Winners, Properly
If you care about snappy first response: 8x A100 + vLLM
The A100 lane was the TTFT winner for both prompt shapes:
- short TTFT: 1.549 s
- long-context TTFT: 0.364 s
And the outputs were coherent, usable English, not quantization soup.
If your user experience is dominated by low-concurrency chat and the first visible token matters most, this is the lane with the best first impression.
If you care about finishing the answer faster and cheaper per token: 8x RTX PRO 6000 + vLLM
The PRO6000 lane won on:
- total latency
- decode throughput
- end-to-end throughput
- generated tokens per USD
Numbers that matter most:
- short decode: 13.11 tok/s vs 6.12 tok/s
- long decode: 22.14 tok/s vs 10.91 tok/s
- long-context tokens/USD: 4,933.65 vs 3,076.72
So if I were renting for actual traffic instead of screenshot-comparison theater, the PRO6000 lane is the better production bet for this exact checkpoint.
The SGLang Plot Twist
This was the annoying part of the run.
Both SGLang lanes did the expensive part correctly:
- environment rebuild finished
- sglang serve launched
- the model loaded fully
- GPUs filled up exactly where you would expect for a fit configuration
And then both lanes failed in the same boring way:
- status stayed at waiting_ready
- port 8000 never appeared
- /v1/models never became reachable
- logs stopped moving after the final shard load
Observed stall signatures:
- 8x RTX PRO 6000 + SGLang 64/64 shards loaded about 76.1 GiB resident per GPU no :8000 listener even after an extended post-load grace window
- 8x A100 + SGLang 64/64 shards loaded about 76.7 GiB resident per GPU no :8000 listener even after an extended post-load grace window
So the honest diagnosis is:
SGLang did not fail because the checkpoint was too large. It failed because the serve stack never crossed the last meter from “loaded model” to “working API.”
So Which Box Would I Rent?
For moonshotai/Kimi-K2.5 on RunPod, today, from this exact setup:
- rent 8x A100 if your priority is lower hourly cost and faster time to first token
- rent 8x RTX PRO 6000 if your priority is better sustained decode and better tokens per dollar
- use vLLM
- do not plan around SGLang in this exact setup unless you are ready to debug a post-load bring-up issue
My own practical pick:
If I were optimizing for user experience in a real service, I would take 8x RTX PRO 6000 + vLLM. The extra $1.60/hr bought enough extra sustained throughput to pay for itself.
If I were optimizing for “first token pops fast and the bill is slightly smaller,” I would take 8x A100 + vLLM.
That is probably the most useful conclusion in the whole report: there are many newer GPUs to choose from, but the best answer still depends on which part of performance you are buying. Startup feel, steady-state throughput, and token economics do not always point to the same machine.
Final Takeaways
- The official native INT4 Kimi checkpoint fit on both 8x A100 and 8x RTX PRO 6000.
- A100 won TTFT on both prompt shapes.
- RTX PRO 6000 won sustained decode speed, total latency, and tokens per USD.
- SGLang on both boxes loaded the model but never exposed a serving endpoint, so this benchmark ended as vLLM: yes, SGLang: not in this stack.
- Token economics are now too important to treat as an afterthought, and the widening menu of GPU options makes setup-specific benchmarking more valuable than broad hardware folklore.
- The simplest summary line is still the best one: A100 won the start, PRO6000 won the race.








