• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Simple math, hard-to-predict LLMs: A post shares Terence Tao’s view

    The post’s account of Tao’s remarks points to natural text’s mix of structure and randomness as a key reason model performance remains difficult to forecast.

    FF
    RP
    4 Sources, ,

    TLDR

    A post summarizing Terence Tao’s remarks says training and running large language models mostly uses linear algebra, matrix multiplication and a bit of calculus—math an undergraduate can handle. The harder puzzle, in this account, is predicting why models succeed at some tasks and fail at others. It points to limited mathematics for data that is partly structured and partly random, like natural text. Without reliable rules for forecasting performance across tasks, it says, progress remains largely empirical.

    Combined views

    540.8K

    4 Sources, first seen 17d ago

    Combined views

    540.8K

    4 Sources, first seen 17d ago

    5.3K likes
    17d ago
    first seen 17d ago
    5.3K likes
    141 comments
    3K saves
    750 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    141 comments
    3K saves
    750 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @rohanpaul_aiTerence Tao: The math behind today’s LLMs is actually simple. Training and running them mostly uses linear algebra, matrix multiplication, and a bit of calculus, material an undergraduate can handle. We understand how to build and operate these models. The real mystery is why they work so well on some tasks and fail on others, and why we cannot predict that in advance. We lack good rules for forecasting performance across tasks, so progress is largely empirical. A key reason is the nature of real-world data. Pure noise is well understood, perfectly structured data is well understood, but natural text sits in between, partly structured and partly random. Mathematics for that middle regime is thin, similar to how physics struggles at meso-scales between atoms and continua. Because of this gap, we can describe the mechanisms but cannot yet explain capability jumps or give reliable task-level predictions. That mismatch, simple machinery versus hard-to-predict behavior, is the core puzzle. ---- Video from Prof @Briankeating YT Channel (Link in comment)
    @francoisfleuretIMO the most reasonable high level model one can have of the AI behemoths is still a brute-force stochastic search in token space where the sampling is biased to mimic mathematician writing that eventually hit target. 1/3

    4 Sources

    @rohanpaul_aiTerence Tao: The math behind today’s LLMs is actually simple. Training and running them mostly uses linear algebra, matrix multiplication, and a bit of calculus, material an undergraduate can handle. We understand how to build and operate these models. The real mystery is why they work so well on some tasks and fail on others, and why we cannot predict that in advance. We lack good rules for forecasting performance across tasks, so progress is largely empirical. A key reason is the nature of real-world data. Pure noise is well understood, perfectly structured data is well understood, but natural text sits in between, partly structured and partly random. Mathematics for that middle regime is thin, similar to how physics struggles at meso-scales between atoms and continua. Because of this gap, we can describe the mechanisms but cannot yet explain capability jumps or give reliable task-level predictions. That mismatch, simple machinery versus hard-to-predict behavior, is the core puzzle. ---- Video from Prof @Briankeating YT Channel (Link in comment)
    @francoisfleuretIMO the most reasonable high level model one can have of the AI behemoths is still a brute-force stochastic search in token space where the sampling is biased to mimic mathematician writing that eventually hit target. 1/3