Prime Intellect launched Prime Inference on Oct. 2, offering serverless endpoints and reserved capacity. Its technical announcement describes a service for frontier open-source models running on the company's GPU infrastructure across multiple datacenters.
The company identifies GLM-5.3 as its first public deployment, which it dates to Sept. 22 on OpenRouter. For developers, the announcement provides a Prime CLI example and an OpenAI-compatible endpoint at https://api.pinference.ai/api/v1.
Separating prompt processing from generation
Prime Intellect describes an architecture that separates the public API from the model fleet, allowing capacity to move or scale without changing the client endpoint. It also describes automatic failover across datacenters to move traffic to healthy deployments.
For GLM-5.3 serving, the company runs prompt processing and token generation on separate GPU groups. NVIDIA Dynamo handles routing and orchestration, while vLLM runs the model. Prime Intellect reports that separating those pools reduced p90 inter-token latency by nearly 40% in its tests, a result specific to its tested serving setup.
The blog also describes retaining conversation history through cached state, including a second cache tier in host memory using Mooncake. The company presents cache reuse as a way to avoid recomputing previously processed history.
Tool calls and planned additions
Prime Intellect describes work on GLM's tool-call format in Dynamo, plus parsing fixes for schema references and nullable argument types. Those changes are part of the company's account of preparing the service for long-running agents.
Its roadmap includes batch and asynchronous inference for offline jobs, plus dedicated and one-click deployments on reserved capacity, including fine-tuned models from Prime training runs.