Wafer reportedly delivers 44% lower latency than Cerebras for YC’s AI Office Hours
Wafer says GLM-5.2 averaged 379 ms on its platform versus 674 ms for Gemma 4 31B on Cerebras—a comparison using different models for YC’s AI Office Hours.
TLDR
Wafer says YC moved its AI Office Hours—offering startup advice from AI versions of YC partners—to a dedicated Wafer endpoint after testing lightweight Gemma and OpenAI models. The company reports 44% lower latency with GLM-5.2 on Wafer than with Gemma 4 31B on Cerebras. It says its agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. Wafer also says users spent 2.5 minutes longer talking to the AI partners on its platform than with other providers.
