Engineer Lists Inference Costs for Large AI Models
Research engineer Florian Brand lists inference prices for four large models of similar scale.
Florian Brand, a research engineer at Prime Intellect, replied with size and price figures for four models. Nemotron is listed at 550B parameters and 2.40/M. Trinity is listed at 398B and 0.80/M. Qwen MoE is listed at 397B and 3.50/M. GLM Flash is listed at 320B and 0.50/M. Brand added that readers should not focus on one parameter of a model. The post supplies no further verification or context beyond these stated values.
Combined views
25.8K
3 posts, first seen 2d ago
Engineer Lists Inference Costs for Large AI Models
Research engineer Florian Brand lists inference prices for four large models of similar scale.
Florian Brand, a research engineer at Prime Intellect, replied with size and price figures for four models. Nemotron is listed at 550B parameters and 2.40/M. Trinity is listed at 398B and 0.80/M. Qwen MoE is listed at 397B and 3.50/M. GLM Flash is listed at 320B and 0.50/M. Brand added that readers should not focus on one parameter of a model. The post supplies no further verification or context beyond these stated values.