Wally announced as an inference stack for open frontier AI models
RunAnywhereAI says Wally is built to be the fastest place to run open frontier models. Its performance snapshot lists GLM-5.3 Max at 790 tokens per second.
TLDR
RunAnywhereAI introduced Wally, its software stack for running open frontier AI models. The company reports these speeds in its performance snapshot: • GLM-5.3 Flash: 380 tokens per second • GLM-5.3 Max: 790 tokens per second • Qwen3.8-27B: 485 tokens per second • DeepSeek-V4.1 Flash: 615 tokens per second
Wally announced as an inference stack for open frontier AI models
RunAnywhereAI says Wally is built to be the fastest place to run open frontier models. Its performance snapshot lists GLM-5.3 Max at 790 tokens per second.