• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Perplexity Details ROSE Model Engine

    ROSE reuses kernels across LLMs and embeddings while skipping KV cache for the latter.

    PE
    1 Source, 26d ago, first seen 26d ago

    TLDR

    Perplexity AI posted on its official X account that ROSE serves as the company's model engine. The engine reuses the same kernels for both LLMs and embeddings. For embeddings it skips the KV cache and applies ragged attention instead of paged attention. ROSE supports multiple attention backends, with kernel selection based on model shape and sequence length. The post included a chart on model throughput.

    Combined views

    1.3K

    1 Source, first seen 26d ago

    Combined views

    1.3K

    1 Source, first seen 26d ago

    13 likes
    13 likes
    1 comments
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    1 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @perplexity_aiROSE is the model engine. It reuses the same kernels for LLMs and embeddings. For embeddings, it skips the KV cache and uses ragged attention instead of paged attention. ROSE supports multiple attention backends, so kernel choice depends on model shape and sequence length.

    1 Source

    @perplexity_aiROSE is the model engine. It reuses the same kernels for LLMs and embeddings. For embeddings, it skips the KV cache and uses ragged attention instead of paged attention. ROSE supports multiple attention backends, so kernel choice depends on model shape and sequence length.