Same concurrency 2500 too Oh boy this makes DS API way more interesting again
It’s here. V4-vision-exp. 117-384 tokens per image, at V4-Flash token prices. I guess it won’t be the best VLM on the market by any measure … except speed and cost. It really is magically fast. In other words, RIP Whale Inference Fleet again.