Generic GPU clusters versus DeepSeek's inference prices
The user points to a large SSD storage cluster as one design detail: caching could fall back to disk, which they say could let cached tokens last 24-plus hours.
TLDR
A user says every GPU cluster they've found has a fairly generic setup, which they suspect reflects a desire to support many workloads without overspecializing. They argue this prevents others from matching DeepSeek's inference prices—the cost of running AI models—and say they'll need to build their own cluster. In a reply, they describe using a large SSD storage cluster so caching can fall back to disk, potentially allowing cached tokens to last 24-plus hours.
Combined views
10.8K
1 Source, first seen 17d ago