• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    AI-as-compiler approach claims speedups over Triton baselines on B200

    A post describes a model that translates source code such as Triton directly into low-level assembly such as PTX, with a verifier checking functional correctness, race conditions and deadlocks.

    AM
    LK
    2 Sources, ,

    TLDR

    The post describes an AI-as-compiler approach that uses a model to translate source code directly into low-level assembly, without manually constructed layers of IR and DSLs. It says a verifier uses symbolic or numeric analysis to check functional correctness and properties such as race conditions and deadlocks. The post reports speedups over Triton baselines on B200 of 1.37x for FlashAttention and 1.10x for Mamba2 state forward.

    Combined views

    18K

    2 Sources, first seen 12h ago

    Combined views

    18K

    2 Sources, first seen 12h ago

    172 likes
    12h ago
    first seen 12h ago
    172 likes
    6 comments
    143 saves
    21 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    6 comments
    143 saves
    21 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @AzaliamirhBringing up the software stack for a new chip has always been a major bottleneck. That's about to change as we enter the era of AI as a compiler! The model directly translates source code (e.g., Triton) to low-level assembly (e.g., PTX), and a verifier checks functional correctness through symbolic / numeric analysis, as well as other properties such as race conditions and deadlocks! No need for manually constructed layers of IR and DSLs. Speedups over Triton baselines on B200: FlashAttention (1.37x) Mamba2 state forward (1.10x) Great work led by @cos_francois, in collaboration with Charly Castes and Thomas Bourgeat from EPFL! More to come soon!
    @lukaszkaiserRT @Azaliamirh: Bringing up the software stack for a new chip has always been a major bottleneck. That's about to change as we enter the er…

    2 Sources

    @AzaliamirhBringing up the software stack for a new chip has always been a major bottleneck. That's about to change as we enter the era of AI as a compiler! The model directly translates source code (e.g., Triton) to low-level assembly (e.g., PTX), and a verifier checks functional correctness through symbolic / numeric analysis, as well as other properties such as race conditions and deadlocks! No need for manually constructed layers of IR and DSLs. Speedups over Triton baselines on B200: FlashAttention (1.37x) Mamba2 state forward (1.10x) Great work led by @cos_francois, in collaboration with Charly Castes and Thomas Bourgeat from EPFL! More to come soon!
    @lukaszkaiserRT @Azaliamirh: Bringing up the software stack for a new chip has always been a major bottleneck. That's about to change as we enter the er…