RL Is Bottlenecked by Inference. Scale It Independently
skypilot.ai
1 thread
what did gpu hours look like here? with 3 replicas for a 1.8x speedup, the cost tradeoff isn’t obvious.
ah, nevermind. 3 engines seem cheaper overall too: 7x661s vs 5x1200s of allocated H100 time per step. Nice.