• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Prime Intellect launches Prime Inference for open models

    Prime Intellect describes separate GPU pools for prompt processing and token generation, plus automatic datacenter failover.

    B(
    WB
    VW
    24 Sources, ,

    TLDR

    Prime Intellect launched Prime Inference with serverless endpoints and reserved capacity for open-source models. Its announcement provides CLI and OpenAI SDK integration, describe separate GPU pools for prompt processing and token generation, and outlines datacenter failover. The company also details tool-call fixes; batch and asynchronous inference and dedicated deployments remain on its roadmap.

    Combined views

    97.9K

    24 Sources, first seen 2h ago

    Combined views

    97.9K

    24 Sources, first seen 2h ago

    1.4K likes
    2h ago
    first seen 2h ago
    1.4K likes
    75 comments
    315 saves
    126 reposts

    Prime Intellect launched Prime Inference on Oct. 2, offering serverless endpoints and reserved capacity. Its technical announcement describes a service for frontier open-source models running on the company's GPU infrastructure across multiple datacenters.

    The company identifies GLM-5.3 as its first public deployment, which it dates to Sept. 22 on OpenRouter. For developers, the announcement provides a Prime CLI example and an OpenAI-compatible endpoint at https://api.pinference.ai/api/v1.

    Separating prompt processing from generation

    Prime Intellect describes an architecture that separates the public API from the model fleet, allowing capacity to move or scale without changing the client endpoint. It also describes automatic failover across datacenters to move traffic to healthy deployments.

    For GLM-5.3 serving, the company runs prompt processing and token generation on separate GPU groups. NVIDIA Dynamo handles routing and orchestration, while vLLM runs the model. Prime Intellect reports that separating those pools reduced p90 inter-token latency by nearly 40% in its tests, a result specific to its tested serving setup.

    The blog also describes retaining conversation history through cached state, including a second cache tier in host memory using Mooncake. The company presents cache reuse as a way to avoid recomputing previously processed history.

    Tool calls and planned additions

    Prime Intellect describes work on GLM's tool-call format in Dynamo, plus parsing fixes for schema references and nullable argument types. Those changes are part of the company's account of preparing the service for long-running agents.

    Its roadmap includes batch and asynchronous inference for offline jobs, plus dedicated and one-click deployments on reserved capacity, including fine-tuned models from Prime training runs.

    Prime Intellect

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    75 comments
    315 saves
    126 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #5

    Today's Rank

    #5

    25 Sources

    primeintellect.aiPrime Inference: Fast, Reliable Serving for Frontier Open Models
    @PrimeIntellectIntroducing Prime Inference: We've served trillions of tokens for RL and dedicated customer deployments To own your intelligence, you need to own your inference Unpacking our inference stack2h
    @chimcisInference is here2h
    @ad0rnai@PrimeIntellect IT'S LIVE2h
    @xeophon🦋 🫶 open source (like @vllm_project and @NVIDIAAI) we will continue to hill climb @michellechen bench and contribute fixes upstream so everyone can benefit2h
    @beffjezosWake up babe new inference provider just dropped2h
    @kevinjosethomasabsolute goats @chimcis @caoshuxun! enabling and encouraging my 10B+ daily token spend 😵‍💫2h
    @kennethnymthe Prime cinematic universe continues to expand ….2h
    @vincentweisserRT @chimcis: Inference is here2h
    @samsja19extremely proud of our inference team, we shifted many of our internal workload to our own glm5 fast endpoint and we are now making it available to everybody2h

    Related

    GolatoTyler says they're joining Prime Intellect as head of scientific AI

    GolatoTyler says they plan to help research teams build scientific agents that learn from cycles of hypotheses, experiments and validated results.

    Vera versus AMD and Intel for RL and agent sandboxes

    A user calls Vera a decent headnode CPU but favors AMD or Intel for reinforcement-learning (RL) and agent sandboxes.

    Prime Intellect releases Prime Sandboxes for reinforcement learning

    Prime Intellect says it built the microVM sandboxes for its own team, citing the complexity and cost of configuring tens of thousands of concurrent sandboxes for model training.

    25 Sources

    primeintellect.aiPrime Inference: Fast, Reliable Serving for Frontier Open Models
    @PrimeIntellectIntroducing Prime Inference: We've served trillions of tokens for RL and dedicated customer deployments To own your intelligence, you need to own your inference Unpacking our inference stack2h
    @chimcisInference is here2h
    @ad0rnai@PrimeIntellect IT'S LIVE2h
    @xeophon🦋 🫶 open source (like @vllm_project and @NVIDIAAI) we will continue to hill climb @michellechen bench and contribute fixes upstream so everyone can benefit2h
    @beffjezosWake up babe new inference provider just dropped2h
    @kevinjosethomasabsolute goats @chimcis @caoshuxun! enabling and encouraging my 10B+ daily token spend 😵‍💫2h
    @kennethnymthe Prime cinematic universe continues to expand ….2h
    @vincentweisserRT @chimcis: Inference is here2h
    @samsja19extremely proud of our inference team, we shifted many of our internal workload to our own glm5 fast endpoint and we are now making it available to everybody2h