ProximalHQ's post-training stack uses Modal GPUs and Kubernetes
A ProximalHQ team member says they dynamically adjust the inference-trainer ratio to maintain training throughput and reduce staleness.
TLDR
A ProximalHQ team member says its post-training stack uses Modal GPUs and Kubernetes. During runs, the team monitors inference throughput and adjusts the inference-trainer ratio to maintain training throughput and reduce staleness. They also use sharded HTTP transfers for weight syncing, saying delta syncing LoRA weights is not worth the decoding overhead. For long-context runs, they say caching only FA4 output and LSE speeds up training steps without running out of memory.
Combined views
8.6K
10 Sources, first seen ago