• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    AI inference needs bandwidth-first hardware, one post argues

    The post highlights Positron, saying it ships Atlas inference machines and is building Asimov, a custom chip that puts model weights next to the multipliers.

    DL
    1 Source, 20d ago, first seen 20d ago

    TLDR

    A September 11 post argues that inference—running a large language model—is limited by memory bandwidth rather than compute. Because model weights are barely reused during the calculations, it proposes designing hardware around streaming those weights, then adding only as much compute as that stream can feed.

    The post points to Positron as an example, describing its Atlas machine and Asimov chip under development. It says the company uses commodity memory instead of high-bandwidth memory and offers an OpenAI-compatible API, making its machines look more like a web service than a programmable GPU.

    Combined views

    2.6K

    1 Source, first seen 20d ago

    Combined views

    2.6K

    1 Source, first seen 20d ago

    16 likes
    16 likes
    2 comments
    10 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    10 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @lemireWhat should be obvious is that inference, running a large language model, is the part that has to be cheap. Most of our hardware was not designed for that. LLM inference is limited by bandwidth. You do a great many matrix-vector multiplications against a huge weight matrix, and you barely reuse those weights. It is closer to streaming a video than to running a simulation or drawing a scene in a game. Graphics processors from Nvidia and others were built for something else. They pileed compute first, then bolted on high-bandwidth memory. It is expensive. An obvious answer is to invert the design. Put the weights first. Make sure you can stream through them at high utilization. Then add only as much compute as the stream can feed. Bandwidth is the scarce resource, not compute. A few companies are working on this. The most interesting American one right now is Positron. They ship an inference machine (Atlas) and are building a custom chip (Asimov) with weights next to the multipliers. Instead of high-bandwidth memory, they use commodity memory. These machines are not programmable, unlike the Nvidia GPUs. They offer an OpenAI-compatible API. So it looks to work more like a web service than a GPU.

    1 Source

    @lemireWhat should be obvious is that inference, running a large language model, is the part that has to be cheap. Most of our hardware was not designed for that. LLM inference is limited by bandwidth. You do a great many matrix-vector multiplications against a huge weight matrix, and you barely reuse those weights. It is closer to streaming a video than to running a simulation or drawing a scene in a game. Graphics processors from Nvidia and others were built for something else. They pileed compute first, then bolted on high-bandwidth memory. It is expensive. An obvious answer is to invert the design. Put the weights first. Make sure you can stream through them at high utilization. Then add only as much compute as the stream can feed. Bandwidth is the scarce resource, not compute. A few companies are working on this. The most interesting American one right now is Positron. They ship an inference machine (Atlas) and are building a custom chip (Asimov) with weights next to the multipliers. Instead of high-bandwidth memory, they use commodity memory. These machines are not programmable, unlike the Nvidia GPUs. They offer an OpenAI-compatible API. So it looks to work more like a web service than a GPU.