Local is no longer the compromise:
- A dense 27B Qwen now beats the 405B Llama from 21 months earlier, on the same RTX 3090 that used to host Llama 2
- Quantize one wrong number, just one, and a model gets 20% dumber. Respect the super weights and GLM 5.2 shrinks from 1.5TB to…