Reaction
Marin is deploying Triton router top-k and XLA scheduling fixes
Marin says the deployment also includes short conv, leaner ragged a2a transport and fused expert backward.
TLDR
Marin says it’s deploying Triton router top-k and short conv, leaner ragged a2a transport, fused expert backward, and XLA scheduling fixes, with details linked on GitHub. In a follow-up, it says the changes mostly address things XLA doesn’t handle automatically, including some unnecessary zeroing, a simple top-k kernel that xla:gpu won’t emit for small arrays, and rematerialization tweaks.
Combined views
22
1 Source, first seen ago