• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Wafer announces AI performance engineering series, starting with transformer inference

    Wafer says the first resource uses worked problems to connect model dimensions to deployment choices, including memory capacity, workload distribution and expected performance under stated assumptions.

    GT
    TA
    CL
    8 Sources, ,

    TLDR

    Wafer says it plans to share every resource from its AI performance engineering repository, starting with “All About Transformer Inference” from How To Scale Your Model. It describes the resource as covering the computations, memory traffic and serving decisions behind running transformer models. Highlighted topics include cache sizing, how batching and quantization shift compute and memory-bandwidth limits, and communication overhead when models are split across hardware.

    Combined views

    475.6K

    8 Sources, first seen 19d ago

    Combined views

    475.6K

    8 Sources, first seen 19d ago

    3.6K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    19d ago
    first seen 19d ago
    3.6K likes
    85 comments
    7.2K saves
    470 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    85 comments
    7.2K saves
    470 reposts

    8 Sources

    @wafer_aiwe launched the most comprehensive ai performance engineering repo in the world last week now we'll be posting every single resource this is Wafer's ai performance engineering series save this as your starting point. links in thread 🧵 part 1: "All About Transformer Inference" from How To Scale Your Model. - the authors cover the computations, memory traffic, and serving decisions behind transformer inference: - arithmetic intensity of linear layers and attention across prefill and decode. - kv cache sizing by layer count, kv heads, head dimension, sequence length, and precision. - the compute/hbm bandwidth crossover and how batch size and quantization shift it. - decode latency and throughput bounds from parameter bytes, kv bytes, and hardware bandwidth. - weight reuse through batching and diminishing throughput gains as kv traffic grows. - gqa, kv quantization, and PagedAttention, including the memory costs each addresses. - model sharding, kv placement, and collective communication overhead. - continuous batching, prefix caching, and disaggregated prefill/decode. the worked problems connect model dimensions to deployment decisions like memory capacity, workload distribution, and expected performance under the stated assumptions.
    @gpusteveyou'll know more about inference than 90% of people if you fully understand this article this is only the first resource in the ai performance engineering repo btw. imagine the ball knowledge in the other ones
    @ycombinatorRT @wafer_ai: we launched the most comprehensive ai performance engineering repo in the world last week now we'll be posting every single…
    @zainhashow inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the inference engine > kv + prefix caching > continuous batching > paged attention > chunked prefill > sampling > agentic loops from inside the engine
    @togethercomputeRT @zainhas: how inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the…
    @garrytanRT @gpusteve: you'll know more about inference than 90% of people if you fully understand this article this is only the first resource in…
    @CShorten30RT @zainhas: how inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the…
    @ChengleiSiRT @zainhas: how inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the…

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    8 Sources

    @wafer_aiwe launched the most comprehensive ai performance engineering repo in the world last week now we'll be posting every single resource this is Wafer's ai performance engineering series save this as your starting point. links in thread 🧵 part 1: "All About Transformer Inference" from How To Scale Your Model. - the authors cover the computations, memory traffic, and serving decisions behind transformer inference: - arithmetic intensity of linear layers and attention across prefill and decode. - kv cache sizing by layer count, kv heads, head dimension, sequence length, and precision. - the compute/hbm bandwidth crossover and how batch size and quantization shift it. - decode latency and throughput bounds from parameter bytes, kv bytes, and hardware bandwidth. - weight reuse through batching and diminishing throughput gains as kv traffic grows. - gqa, kv quantization, and PagedAttention, including the memory costs each addresses. - model sharding, kv placement, and collective communication overhead. - continuous batching, prefix caching, and disaggregated prefill/decode. the worked problems connect model dimensions to deployment decisions like memory capacity, workload distribution, and expected performance under the stated assumptions.
    @gpusteveyou'll know more about inference than 90% of people if you fully understand this article this is only the first resource in the ai performance engineering repo btw. imagine the ball knowledge in the other ones
    @ycombinatorRT @wafer_ai: we launched the most comprehensive ai performance engineering repo in the world last week now we'll be posting every single…
    @zainhashow inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the inference engine > kv + prefix caching > continuous batching > paged attention > chunked prefill > sampling > agentic loops from inside the engine
    @togethercomputeRT @zainhas: how inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the…
    @garrytanRT @gpusteve: you'll know more about inference than 90% of people if you fully understand this article this is only the first resource in…
    @CShorten30RT @zainhas: how inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the…
    @ChengleiSiRT @zainhas: how inference engines actually work - releasing full talk slides! i cover everything in the lifetime of a request e2e: > the…