GLM 5.3 Flash Runs Untuned on Dual DGX Stations
Alec Fong shares untuned inference results for the model on his hardware setup.
Alec Fong posted that GLM 5.3 Flash will likely serve as his team's new daily driver. He ran the model untuned on two DGX stations with tensor parallelism of 2. The setup delivered 881 tokens per second at context length 64 and 232 tokens per second at context length 1. Fong noted the model supports vision and could handle four to eight simultaneous users at longer contexts. A retweet of the post pointed to remaining opportunities in prediction heads and unoptimized kernels by comparing the machine's cost performance to recent discounts on smaller systems.
Combined views
91.7K
3 posts, first seen 15h ago