Announcement
Photon 2.6 adds FP8 inference and speculative decoding; Qwen3.5-27B is claimed to run at over 400 tokens/sec on one B200
Moondream says FP8 and DFlash speculative verification are now part of its megakernel compiler.
TLDR
A post announcing Photon 2.6 says it supports FP8 inference and speculative decoding. Moondream says Photon runs Qwen3.5-27B at over 400 tokens per second on one B200, and that FP8 and DFlash speculative verification are now part of its megakernel compiler.
Combined views
3.4K
2 Sources, first seen 1h ago
likes
