Perplexity Details ROSE Model Engine
ROSE reuses kernels across LLMs and embeddings while skipping KV cache for the latter.
TLDR
Perplexity AI posted on its official X account that ROSE serves as the company's model engine. The engine reuses the same kernels for both LLMs and embeddings. For embeddings it skips the KV cache and applies ragged attention instead of paged attention. ROSE supports multiple attention backends, with kernel selection based on model shape and sequence length. The post included a chart on model throughput.
Combined views
1.3K
1 Source, first seen 26d ago