• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Jev-compatible public API reportedly runs 64 parallel generation tasks in under a second

    Its creator says the API runs the open model Qwen3.6-35B-A3B and uses SGLang’s radix cache to reuse initial prompt processing for parallel generation.

    Nando de FreitasND
    Susan ZhangSZ
    Eric ZhangEZ
    8 Sources, ,

    TLDR

    The public API, openjev-sglang, is an experiment in how well an off-the-shelf model works, its creator says. They report 64 parallel generation tasks in under a second using Qwen3.6-35B-A3B and SGLang’s radix cache. The GitHub project describes the endpoint as prefill-only. In a follow-up reply, the creator shared an endpoint for users to connect their frontends.

    Combined views

    163K

    8 Sources, first seen 20d ago

    Combined views

    163K

    8 Sources, first seen 20d ago

    1.2K likes
    20d ago
    first seen 20d ago
    1.2K likes
    40 comments
    968 saves
    85 reposts
    40 comments
    968 saves
    85 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    8 Sources

    Eric Zhang@ekzhang1Inspired by @typesafeai , here is a Jev-compatible public API to play with It runs a comparable open model (Qwen3.6-35B-A3B), and just uses SGLang radix cache to preserve the prefill reuse / really fast parallel systemone generation - 64 tasks in <1s. https://github.com/ekzhang/openjev-sglang20d
    Susan Zhang@suchenzangTechnically, the most this could mean is for the calibration data used for training to skew towards identifying as Qwen.19d
    Nando de Freitas@NandoDFRT @Zefan_Cai: Inspired by Jev, we built Open-Jev: open-source decision models. 2B/9B: LoRA adapters + decision heads. Code: https://t.co/…17d

    8 Sources

    Eric Zhang@ekzhang1Inspired by @typesafeai , here is a Jev-compatible public API to play with It runs a comparable open model (Qwen3.6-35B-A3B), and just uses SGLang radix cache to preserve the prefill reuse / really fast parallel systemone generation - 64 tasks in <1s. https://github.com/ekzhang/openjev-sglang20d
    Susan Zhang@suchenzangTechnically, the most this could mean is for the calibration data used for training to skew towards identifying as Qwen.19d
    Nando de Freitas@NandoDFRT @Zefan_Cai: Inspired by Jev, we built Open-Jev: open-source decision models. 2B/9B: LoRA adapters + decision heads. Code: https://t.co/…17d