RL Is Bottlenecked by Inference. Scale It Independently
1–4 of 4 posts
Re: RL Is Bottlenecked by Inference. Scale It Independently
#2[deleted]
Re: RL Is Bottlenecked by Inference. Scale It Independently
#3what did gpu hours look like here? with 3 replicas for a 1.8x speedup, the cost tradeoff isn’t obvious.
Re: RL Is Bottlenecked by Inference. Scale It Independently
#4what did gpu hours look like here? with 3 replicas for a 1.8x speedup, the cost tradeoff isn’t obvious.
ah, nevermind. 3 engines seem cheaper overall too: 7x661s vs 5x1200s of allocated H100 time per step. Nice.