Researchers Consider Cached Tokens in LLM Decode Costs
You Jiacheng and Horace He examine token counts for LLM inference billing.
You Jiacheng posted that every decode step reloads all cached tokens from the KV cache. Horace He replied that counting every such token would be a pretty fair way to bill. Both work on AI and machine learning topics. You Jiacheng is a PhD student at Tsinghua IIIS. Horace He is a research engineer formerly at Meta and now at Thinking Machines Lab. The exchange questions standard token usage metrics for large language model inference.
Combined views
5.1K
2 posts, first seen 9h ago
Researchers Consider Cached Tokens in LLM Decode Costs
You Jiacheng and Horace He examine token counts for LLM inference billing.
You Jiacheng posted that every decode step reloads all cached tokens from the KV cache. Horace He replied that counting every such token would be a pretty fair way to bill. Both work on AI and machine learning topics. You Jiacheng is a PhD student at Tsinghua IIIS. Horace He is a research engineer formerly at Meta and now at Thinking Machines Lab. The exchange questions standard token usage metrics for large language model inference.


