• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Running large language models is a bandwidth problem, one post argues

    A post points to Positron, saying it ships an Atlas inference machine and is building an Asimov chip with model weights next to the multipliers, using commodity memory rather than high-bandwidth memory.

    PD
    DJ
    2 Sources, 19d ago, first seen 19d ago

    TLDR

    A post argues that inference—running a large language model—is limited by bandwidth: it requires reading huge matrices of model weights with little reuse. The author advocates designing hardware around streaming those weights, then adding only as much compute as that stream can feed. The post highlights Positron, saying it ships the Atlas inference machine and is developing Asimov, a custom chip with weights next to multipliers. It says the machines use commodity memory instead of high-bandwidth memory and are not programmable like Nvidia GPUs, instead offering an OpenAI-compatible API.

    Combined views

    14.9K

    2 Sources, first seen 19d ago

    Combined views

    14.9K

    2 Sources, first seen 19d ago

    227 likes
    227 likes
    43 comments
    18 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    43 comments
    18 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @zehavocRT @lemire: What should be obvious is that inference, running a large language model, is the part that has to be cheap. Most of our hardwar…
    @pmddomingosCurrent AI computing is so extraordinarily inefficient that Nvidia won't have much of a moat when a better solution appears.

    2 Sources

    @zehavocRT @lemire: What should be obvious is that inference, running a large language model, is the part that has to be cheap. Most of our hardwar…
    @pmddomingosCurrent AI computing is so extraordinarily inefficient that Nvidia won't have much of a moat when a better solution appears.