• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Sebastian Raschka Shares LLM Text Generation Video

    Video covers KV caching to prepare base models for later reasoning techniques.

    SR
    1 Source, 24d ago, first seen 24d ago

    TLDR

    Sebastian Raschka posted a video on X explaining the text generation process in LLMs along with KV caching. The post lists timestamps for an introduction and demo, book guidance, a Chapter 2 overview, and PyTorch checks. It forms round 2 of his reasoning from scratch series and readies the base model ahead of upcoming reasoning techniques. The attached media is a screen recording.

    Combined views

    52K

    1 Source, first seen 24d ago

    Combined views

    52K

    1 Source, first seen 24d ago

    1.2K likes
    1.2K likes
    43 comments
    1.1K saves
    143 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    43 comments
    1.1K saves
    143 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @rasbtReasoning from scratch round 2: In this video, I cover the text generation process in LLMs and KV caching (to prepare the base model before adding reasoning techniques in the upcoming ones). 00:00 Introduction and reasoning model demo 01:55 How to work through the book 05:00 Chapter 2 overview 08:25 Checking PyTorch and hardware support 10:26 Apple silicon and MPS caveats 15:00 Cloud GPU options 16:08 Tokens and tokenization 18:20 Qwen3 and the Reasoning From Scratch package 23:05 Encoding and decoding text 26:24 Downloading weights and selecting a device 31:01 Loading the pretrained Qwen3 model 34:32 How LLMs generate text 36:47 Input tensors and batch dimensions 41:48 Running the model in inference mode 44:11 Logits and next-token predictions 49:21 Greedy decoding with argmax 52:28 Building a streaming text generator 01:01:28 Generating text and handling end-of-sequence tokens 01:06:00 Benchmarking text generation 01:14:34 How KV caching works 01:17:22 Adding KV caching and measuring the speedup 01:24:31 Model compilation with torch.compile 01:30:33 Combining compilation with KV caching 01:32:53 Comparing CPU and GPU performance 01:35:32 Recap and next steps

    1 Source

    @rasbtReasoning from scratch round 2: In this video, I cover the text generation process in LLMs and KV caching (to prepare the base model before adding reasoning techniques in the upcoming ones). 00:00 Introduction and reasoning model demo 01:55 How to work through the book 05:00 Chapter 2 overview 08:25 Checking PyTorch and hardware support 10:26 Apple silicon and MPS caveats 15:00 Cloud GPU options 16:08 Tokens and tokenization 18:20 Qwen3 and the Reasoning From Scratch package 23:05 Encoding and decoding text 26:24 Downloading weights and selecting a device 31:01 Loading the pretrained Qwen3 model 34:32 How LLMs generate text 36:47 Input tensors and batch dimensions 41:48 Running the model in inference mode 44:11 Logits and next-token predictions 49:21 Greedy decoding with argmax 52:28 Building a streaming text generator 01:01:28 Generating text and handling end-of-sequence tokens 01:06:00 Benchmarking text generation 01:14:34 How KV caching works 01:17:22 Adding KV caching and measuring the speedup 01:24:31 Model compilation with torch.compile 01:30:33 Combining compilation with KV caching 01:32:53 Comparing CPU and GPU performance 01:35:32 Recap and next steps