• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Perplexity Open-Sources Lily Local Inference Engine

    The engine powers hybrid compute tasks in the company's Mac app.

    AS
    B(
    YS
    9 Sources, 28d ago, first seen 28d ago

    TLDR

    Perplexity announced it is open-sourcing Lily, a local inference engine built for Apple silicon. The release targets the Qwen3.6-35B-A3B model and supports hybrid compute in the Mac app by keeping on-device work from slowing cloud tasks. Company statements describe custom Metal kernels and a compact Rust runtime that manages session state and generation. Posts from Perplexity and its executives note the engine was purpose-built rather than adapted from general frameworks. The code appears in the pplx-garden repository, and a company blog post covers the on-device optimizations.

    Combined views

    625.1K

    9 Sources, first seen 28d ago

    Combined views

    625.1K

    9 Sources, first seen 28d ago

    5.4K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5.4K likes
    197 comments
    2.8K saves
    486 reposts
    197 comments
    2.8K saves
    486 reposts

    Sentiment

    Positive88.9%11.1%Negative

    Summary

    Sentiment

    Positive88.9%11.1%Negative

    Many accounts welcomed Perplexity open-sourcing Lily, its local inference engine for Apple Silicon optimized on Qwen MoE, due to strong benchmark speedups, while some criticized the M5+ hardware limits and unrealistic test setup.

    Based on 47 sentiment-bearing replies from 45 accounts across 3 conversations.

    Summary

    Many accounts welcomed Perplexity open-sourcing Lily, its local inference engine for Apple Silicon optimized on Qwen MoE, due to strong benchmark speedups, while some criticized the M5+ hardware limits and unrealistic test setup.

    Based on 47 sentiment-bearing replies from 45 accounts across 3 conversations.

    9 Sources

    @perplexity_aiToday we’re open-sourcing Lily, the local inference engine we built for hybrid compute in Perplexity Computer. Lily is specialized for Qwen3.6-35B-A3B on Apple silicon, built so on-device compute doesn’t bottleneck Computer tasks. Read more: https://www.perplexity.ai/hub/blog/optimizing-on-device-inference-for-apple-silicon
    @denisyaratswe are open-sourcing Lily, lightweight local inference engine for Apple Silicon, optimized directly for the Qwen MoE architecture and implemented on top of custom Metal kernels and a compact Rust runtime that manages both the session state and the generation loop. we carefully optimized the engine around Qwen's exact architecture and dimensions: matrix/vector kernels for prefill/decode, minimal data movement across expert, attention, and recurrent paths, and tiles and layouts tuned to the live workload. with these optimizations we achieve 1.23x faster prefill and 1.35x faster decode than MLX-LM on Qwen3.6-35B-A3B Q4! for more details, check out our blog and code: code: https://github.com/perplexityai/pplx-garden/tree/main/lily blog: https://www.perplexity.ai/hub/blog/optimizing-on-device-inference-for-apple-silicon
    @AravSrinivasWe're open-sourcing Lily, Perplexity's local inference engine for serving models locally on Apple Silicon. This powers Perplexity's newly introduced hybrid compute feature for the Mac app.
    @beffjezosPerplexity going all-in on on-device inference. Very interesting
    @ying11231RT @AravSrinivas: We're open-sourcing Lily, Perplexity's local inference engine for serving models locally on Apple Silicon. This powers Pe…

    9 Sources

    @perplexity_aiToday we’re open-sourcing Lily, the local inference engine we built for hybrid compute in Perplexity Computer. Lily is specialized for Qwen3.6-35B-A3B on Apple silicon, built so on-device compute doesn’t bottleneck Computer tasks. Read more: https://www.perplexity.ai/hub/blog/optimizing-on-device-inference-for-apple-silicon
    @denisyaratswe are open-sourcing Lily, lightweight local inference engine for Apple Silicon, optimized directly for the Qwen MoE architecture and implemented on top of custom Metal kernels and a compact Rust runtime that manages both the session state and the generation loop. we carefully optimized the engine around Qwen's exact architecture and dimensions: matrix/vector kernels for prefill/decode, minimal data movement across expert, attention, and recurrent paths, and tiles and layouts tuned to the live workload. with these optimizations we achieve 1.23x faster prefill and 1.35x faster decode than MLX-LM on Qwen3.6-35B-A3B Q4! for more details, check out our blog and code: code: https://github.com/perplexityai/pplx-garden/tree/main/lily blog: https://www.perplexity.ai/hub/blog/optimizing-on-device-inference-for-apple-silicon
    @AravSrinivasWe're open-sourcing Lily, Perplexity's local inference engine for serving models locally on Apple Silicon. This powers Perplexity's newly introduced hybrid compute feature for the Mac app.
    @beffjezosPerplexity going all-in on on-device inference. Very interesting
    @ying11231RT @AravSrinivas: We're open-sourcing Lily, Perplexity's local inference engine for serving models locally on Apple Silicon. This powers Pe…