Many users celebrated NVIDIA Blackwell Ultra's record DeepSeek-V3 training speeds as a win for open-source AI progress, while others criticized the results for potentially aiding competitors or skirting export rules.
Based on 27 visible X reactions from 49 accounts; directional sample.
Ask a question below.
Published answers will appear here.
The GB300 platform tripled the throughput of GB200 hardware.
@NVIDIAAI I'm not saying that NVIDIA engaged in a conspiracy to violate export control rules... but this really sounds like bragging about evading the export control rules for the GB300. cc: @howardlutnick @mkratsios47 @UnderSecE @SecScottBessent, what's the deal? https://x.com/NVIDIAAI/status/2079582939373863353
@NVIDIAAI Mind-blowing performance! Crushing records like this is exactly why NVIDIA continues to lead the AI resolution
@teortaxesTex Just how much longer is US tech going to embarass itself by making NVIDIA do their reference materials on DeepSeek?
@NVIDIAAI Nvidia Blackwell Ultra is a massive win for open-source AI. 🚀
@NVIDIAAI Take my money please.
1648 TFLOPS/GPU on GB300 sounds really cool, and is. this is with MXFP8. So 33% MFU. DeepSeek themselves, in late 2023, did ≈25% on H800s (they calculate MFU based on BF16 here, but heavily used FP8). A 4.3x gap in per-GPU speed. About 2.5x in Watt/tokens? But I expected more.
NVIDIA Blackwell Ultra just set a world record for pre-training DeepSeek-V3 671B, achieving 1,648 TFLOPs per GPU. This is about 3x the earlier delivered performance of the previous generation, and is a result of our extreme co-design and continuous software optimization across popular frameworks including Megatron-Core, TorchTitan, and JAX.
Many users celebrated NVIDIA Blackwell Ultra's record DeepSeek-V3 training speeds as a win for open-source AI progress, while others criticized the results for potentially aiding competitors or skirting export rules.
Based on 27 visible X reactions from 49 accounts; directional sample.
Ask a question below.
Published answers will appear here.
Just did the mafs and it roughly checks out btw 0,385*2048*3600*24*55 = 3.74e24 (PFLOPs*H800s*training run duration) 37e9*14.8e12*6 = 3.28e24 I guess the discrepancy is due to the bf16 proportion 50K hoppers bros in shambles
Can’t stap/won’t stap
Really cool how V3 is still the standard benchmark