• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    @vikhyatk Describes Photon 2.0 Megakernel Compiler

    Founder Vikhyat K. describes the compiler that fuses full forward passes into single GPU programs.

    SZ
    SD
    KA
    18 Sources, 58d ago, first seen 58d ago

    TLDR

    Vikhyat K., co-founder and CTO of Moondream AI, posted that the team grew tired of hand-tuning GPU kernels and built Photon 2.0 instead. The compiler converts the complete forward pass of Moondream, Qwen 3.5, and Gemma 4 into megakernels that run as single GPU programs. He illustrated the approach with a tracer DSL example that maps data flows and dependencies for a dense Moondream layer. Other posts in the thread include images showing resulting GPU behavior.

    Combined views

    339.2K

    18 Sources, first seen 58d ago

    Combined views

    339.2K

    18 Sources, first seen 58d ago

    2.6K likes
    2.6K likes
    88 comments
    1.3K saves
    359 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    88 comments
    1.3K saves
    359 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    18 Sources

    @vikhyatkGot sick of hand-tuning GPU kernels, so we built a compiler. Photon 2.0 compiles Moondream, Qwen 3.5, and Gemma 4 into megakernels: the entire forward pass in one GPU program.
    @suchenzangthe one path to data-efficiency is to abuse GPUs like they've never been abused before
    @charles_irllit
    @DanielleFongRT @alexinexxx: great launch, just had to make one small adjustment
    @andersonbcdefg1) what
    @xlr8harderRT @vikhyatk: Got sick of hand-tuning GPU kernels, so we built a compiler. Photon 2.0 compiles Moondream, Qwen 3.5, and Gemma 4 into megak…
    @xeophon@vikhyatk Banger banger banger
    @yacineMTBRT @vikhyatk: The compiler then explores various ways to schedule the work on the GPU SMs. Actually compiling the megakernel is slow, so we…
    @sedielemRT @vikhyatk: Got sick of hand-tuning GPU kernels, so we built a compiler. Photon 2.0 compiles Moondream, Qwen 3.5, and Gemma 4 into megak…

    18 Sources

    @vikhyatkGot sick of hand-tuning GPU kernels, so we built a compiler. Photon 2.0 compiles Moondream, Qwen 3.5, and Gemma 4 into megakernels: the entire forward pass in one GPU program.
    @suchenzangthe one path to data-efficiency is to abuse GPUs like they've never been abused before
    @charles_irllit
    @DanielleFongRT @alexinexxx: great launch, just had to make one small adjustment
    @andersonbcdefg1) what
    @xlr8harderRT @vikhyatk: Got sick of hand-tuning GPU kernels, so we built a compiler. Photon 2.0 compiles Moondream, Qwen 3.5, and Gemma 4 into megak…
    @xeophon@vikhyatk Banger banger banger
    @yacineMTBRT @vikhyatk: The compiler then explores various ways to schedule the work on the GPU SMs. Actually compiling the megakernel is slow, so we…
    @sedielemRT @vikhyatk: Got sick of hand-tuning GPU kernels, so we built a compiler. Photon 2.0 compiles Moondream, Qwen 3.5, and Gemma 4 into megak…