• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI Judge Spots FlashInfer Edge Cases in Attention

    Josh Tobin says an AI judge his team built spotted the issues during new autoresearch work.

    JT
    1 Source, 34d ago, first seen 34d ago

    TLDR

    Josh Tobin, co-founder of Recursive Superintelligence, posted that his group's automated research tool found edge cases in FlashInfer that could affect inference performance in vLLM and SGLang. The team had previously built a reward hacking judge for performance optimization tasks. Tobin stated the judge identified the problems while running on fresh autoresearch experiments and that the group helped produce a fix. The post links to an arXiv paper on Flash Attention stability and a FlashInfer GitHub pull request addressing extreme negative logits in masked softmax.

    Combined views

    80.2K

    1 Source, first seen 34d ago

    Combined views

    80.2K

    1 Source, first seen 34d ago

    400 likes
    400 likes
    12 comments
    328 saves
    42 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    12 comments
    328 saves
    42 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @josh_tobin_A fun small win for automated research: We found and helped fix some edge cases that could affect inference performance in vLLM and SGLang. In our last blog post, we described a reward hacking judge we developed for performance optimization tasks. While applying the judge to some new autoresearch work, it discovered that some FlashInfer (a library underpinning vLLM and SGLang) kernels used a hard coded value of -50,000 as a masked-attention sentinel -- even though valid QK values can be smaller. Corner cases like this that are numerically wrong but silent can cause a huge amount of headache to find and fix, e.g., the historical debate around flash attention (https://arxiv.org/abs/2405.02803). Great use case for AI. https://github.com/flashinfer-ai/flashinfer/pull/4401

    1 Source

    @josh_tobin_A fun small win for automated research: We found and helped fix some edge cases that could affect inference performance in vLLM and SGLang. In our last blog post, we described a reward hacking judge we developed for performance optimization tasks. While applying the judge to some new autoresearch work, it discovered that some FlashInfer (a library underpinning vLLM and SGLang) kernels used a hard coded value of -50,000 as a masked-attention sentinel -- even though valid QK values can be smaller. Corner cases like this that are numerically wrong but silent can cause a huge amount of headache to find and fix, e.g., the historical debate around flash attention (https://arxiv.org/abs/2405.02803). Great use case for AI. https://github.com/flashinfer-ai/flashinfer/pull/4401