Running large language models is a bandwidth problem, one post argues
A post points to Positron, saying it ships an Atlas inference machine and is building an Asimov chip with model weights next to the multipliers, using commodity memory rather than high-bandwidth memory.
TLDR
A post argues that inference—running a large language model—is limited by bandwidth: it requires reading huge matrices of model weights with little reuse. The author advocates designing hardware around streaming those weights, then adding only as much compute as that stream can feed. The post highlights Positron, saying it ships the Atlas inference machine and is developing Asimov, a custom chip with weights next to multipliers. It says the machines use commodity memory instead of high-bandwidth memory and are not programmable like Nvidia GPUs, instead offering an OpenAI-compatible API.
Combined views
14.9K
2 Sources, first seen 19d ago