• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Perplexity Publishes Research on Fast GPU Embeddings

    Perplexity details its GPU infrastructure for embedding models used in search ranking.

    AS
    PE
    DY
    4 Sources, 26d ago, first seen 26d ago

    TLDR

    Perplexity AI posted research on its serving infrastructure behind embedding and ranking models. The company described components including Ivy, Tulip, and ROSE for handling batching, latency, and throughput on GPUs. Founder Aravind Srinivas called it a deep dive into how Perplexity serves search results at scale: embeddings for ranking, GPU-based model inference, request batching, running inference servers, and handling latency/throughput trade-offs. The post links to a blog post titled Fast Embeddings on GPUs. Replies on X mixed praise for the technical details with criticism of the product.

    Combined views

    277.4K

    4 Sources, first seen 26d ago

    Combined views

    277.4K

    4 Sources, first seen 26d ago

    1.6K likes
    1.6K likes
    92 comments
    771 saves
    82 reposts
    92 comments
    771 saves
    82 reposts

    Sentiment

    Positive84.2%15.8%Negative

    Summary

    Sentiment

    Positive84.2%15.8%Negative

    Many accounts praised Perplexity’s engineering for serving embeddings and rerankers at exabyte scale and showed interest in the related inference and infrastructure roles, while a few called the system untrustworthy or useless.

    Based on 20 sentiment-bearing replies from 19 accounts across 2 conversations.

    Summary

    Many accounts praised Perplexity’s engineering for serving embeddings and rerankers at exabyte scale and showed interest in the related inference and infrastructure roles, while a few called the system untrustworthy or useless.

    Based on 20 sentiment-bearing replies from 19 accounts across 2 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @perplexity_aiEvery answer in Perplexity starts with embedding and ranking models picking the most relevant results for the query. Today we published research on how we built SoTA serving infrastructure behind those models. Read the research: https://www.perplexity.ai/hub/blog/fast-embeddings-on-gpus
    @AravSrinivasA deep dive into how Perplexity serves search results at scale: embeddings for ranking, GPU-based model inference, request batching, running inference servers, and handling latency/throughput trade-offs.
    @denisyaratscheck out our new blog post on how we serve embeddings and rerankers for our SOTA pplx-embed models over an exabyte-scale search index. embedding models have a similar serving profile to LLMs: batch indexing is compute-bound prefill, online serving is memory-bound decode. this lets us reuse the same optimized kernels we built for large LLMs and get great efficiency for free. on top of that, we optimized the runtime: whole-model CUDA graphs captured lazily as the engine serves, and a LazyTensor in Rust that overlaps CPU scheduling with GPU execution. result: up to 3x lower p50 and 4.8x lower p99 latency than vLLM on BGE-M3 at 128 tokens, single H200! if you are interested in working on problems like this, DM me or apply at http://perplexity.ai/hub/careers

    4 Sources

    @perplexity_aiEvery answer in Perplexity starts with embedding and ranking models picking the most relevant results for the query. Today we published research on how we built SoTA serving infrastructure behind those models. Read the research: https://www.perplexity.ai/hub/blog/fast-embeddings-on-gpus
    @AravSrinivasA deep dive into how Perplexity serves search results at scale: embeddings for ranking, GPU-based model inference, request batching, running inference servers, and handling latency/throughput trade-offs.
    @denisyaratscheck out our new blog post on how we serve embeddings and rerankers for our SOTA pplx-embed models over an exabyte-scale search index. embedding models have a similar serving profile to LLMs: batch indexing is compute-bound prefill, online serving is memory-bound decode. this lets us reuse the same optimized kernels we built for large LLMs and get great efficiency for free. on top of that, we optimized the runtime: whole-model CUDA graphs captured lazily as the engine serves, and a LazyTensor in Rust that overlaps CPU scheduling with GPU execution. result: up to 3x lower p50 and 4.8x lower p99 latency than vLLM on BGE-M3 at 128 tokens, single H200! if you are interested in working on problems like this, DM me or apply at http://perplexity.ai/hub/careers