Jev-compatible public API reportedly runs 64 parallel generation tasks in under a second
Its creator says the API runs the open model Qwen3.6-35B-A3B and uses SGLang’s radix cache to reuse initial prompt processing for parallel generation.
TLDR
The public API, openjev-sglang, is an experiment in how well an off-the-shelf model works, its creator says. They report 64 parallel generation tasks in under a second using Qwen3.6-35B-A3B and SGLang’s radix cache. The GitHub project describes the endpoint as prefill-only. In a follow-up reply, the creator shared an endpoint for users to connect their frontends.
Combined views
23.8K
3 Sources, first seen 9h ago
Jev-compatible public API reportedly runs 64 parallel generation tasks in under a second
Its creator says the API runs the open model Qwen3.6-35B-A3B and uses SGLang’s radix cache to reuse initial prompt processing for parallel generation.
TLDR
The public API, openjev-sglang, is an experiment in how well an off-the-shelf model works, its creator says. They report 64 parallel generation tasks in under a second using Qwen3.6-35B-A3B and SGLang’s radix cache. The GitHub project describes the endpoint as prefill-only. In a follow-up reply, the creator shared an endpoint for users to connect their frontends.

