Slow calls during bursts and AI agent deployment decisions
One slow call can stall an agent's parallel work, the user says. They favor p99 latency—the slow end of response times—over median tokens per second as a deployment metric.
TLDR
A reply argues that agent deployments hinge on p99 latency, or 99th-percentile response time, during bursts—not median tokens per second. The user says agents make calls in parallel, so one slow call can stall the whole loop. They also point to cache rate as helpful for the same reason.
Combined views
8
1 Source, first seen 18d ago
reposts