Report
TokenRouter reportedly boosts throughput for per-token routing between small and large AI models
A post sharing the Tsinghua paper reports 2.01–64.15× the throughput of a stronger existing setup across five routing methods.
TLDR
A post sharing a new Tsinghua paper says TokenRouter gives small and large models separate servers and passes partly written answers between them while retaining the KV cache—the models’ memory of the text so far. It briefly holds requests so each model can work on larger batches. Across five routing methods, the post reports 2.01–64.15× the throughput of a stronger existing setup.
Combined views
—
1 Source, first seen ago
— likes— comments— saves— reposts
