RL Is Bottlenecked by Inference. Scale It Independently

10 pointsposted 4 hours ago
by alex000kim

3 Comments

user

4 hours ago

[deleted]

efiop

4 hours ago

what did gpu hours look like here? with 3 replicas for a 1.8x speedup, the cost tradeoff isn’t obvious.

efiop

4 hours ago

ah, nevermind. 3 engines seem cheaper overall too: 7x661s vs 5x1200s of allocated H100 time per step. Nice.