• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Slow calls during bursts and AI agent deployment decisions

    One slow call can stall an agent's parallel work, the user says. They favor p99 latency—the slow end of response times—over median tokens per second as a deployment metric.

    VV
    1 Source, 18d ago, first seen 18d ago

    TLDR

    A reply argues that agent deployments hinge on p99 latency, or 99th-percentile response time, during bursts—not median tokens per second. The user says agents make calls in parallel, so one slow call can stall the whole loop. They also point to cache rate as helpful for the same reason.

    Combined views

    8

    1 Source, first seen 18d ago

    reposts

    Combined views

    8

    1 Source, first seen 18d ago

    1 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @vipulvedRT @vikasmalpani: @vipulved @OpenRouter @togethercompute Right, and the dimension that quietly decides agent deployments is p99 under a bur…

    1 Source

    @vipulvedRT @vikasmalpani: @vipulved @OpenRouter @togethercompute Right, and the dimension that quietly decides agent deployments is p99 under a bur…