Horace He Calls Cached Token Counts Incredibly Dumb
Researchers note that visible dashboard numbers lead users to include cached inputs when reporting token usage.
TLDR
Horace He posted on X that people discussing LLM token usage usually include cached input tokens and called the practice incredibly dumb. Hieu Pham replied that users read the most visible dashboard number rather than seeking a breakdown. Lucas Beyer joked that a "kv tokens" metric multiplying by layer count would make the numbers even bigger. Horace He followed up suggesting counting every decode step load of cached tokens. The posts discuss reported LLM token counts that include cached inputs.
Horace He Calls Cached Token Counts Incredibly Dumb
Researchers note that visible dashboard numbers lead users to include cached inputs when reporting token usage.