• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Reaction

Marin is deploying Triton router top-k and XLA scheduling fixes

Marin says the deployment also includes short conv, leaner ragged a2a transport and fused expert backward.

David HallDH
1 Source, 26m ago, first seen 26m ago

TLDR

Marin says it’s deploying Triton router top-k and short conv, leaner ragged a2a transport, fused expert backward, and XLA scheduling fixes, with details linked on GitHub. In a follow-up, it says the changes mostly address things XLA doesn’t handle automatically, including some unnecessary zeroing, a simple top-k kernel that xla:gpu won’t emit for small arrays, and rematerialization tweaks.

Combined views

22

1 Source, first seen 26m ago

Combined views

22

1 Source, first seen 26m ago

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

David Hall@dlwhThey mostly fall into the "XLA is very good, but it's not magic" category: avoiding some unnecessary zeroing, a pretty simple topk kernel (which xla:gpu won't emit for small arrays), tweaking rematerialization, etc26m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    David Hall@dlwhThey mostly fall into the "XLA is very good, but it's not magic" category: avoiding some unnecessary zeroing, a pretty simple topk kernel (which xla:gpu won't emit for small arrays), tweaking rematerialization, etc26m
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet