Based on 23 visible X reactions from 38 accounts; directional sample.
Ask a question below.
Published answers will appear here.
Physical hardware testing recorded 800,000 tokens per second per megawatt.
@thdxr I mean, it's exactly what Nvidia said it was when they announced in March. https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform You got me all excited that was it was a new surprising thing
@thdxr Yesterday I read how they used 8,000 nodes and trained a large deepseek model in 2 minutes. Absolutely nutty.
@thdxr This is insane. Cost per token is only going to continue to become cheaper.
@thdxr it’s true, and we still early stage.
Again, I expected more.
@thdxr I want an RTX 7000:D
The underrated thing is that they're running R1 at ≈150 tps, about 7 times faster than it was served by DeepSeek proper. And for Rubin, this is a healthy regime, not desperate at all, 60% of peak throughput. Almost 1 t/s/watt. V4 does roughly this, but with DSpark and 2x slower.
Based on 23 visible X reactions from 38 accounts; directional sample.
Ask a question below.
Published answers will appear here.
what the hell no way