• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Claude Opus 5 reportedly gets all 1,200 multiplication problems right without tools at maximum reasoning

    A user says the test covered numbers up to 20 digits each. Their follow-up reports roughly 0% accuracy for Fable 5.1 on 5-digit-by-6-digit multiplication at low reasoning.

    j⧉nusJ⧉
    Pasquale MinerviniPM
    elieEL
    18 Sources, ,

    TLDR

    A user reported on September 16 that Claude Opus 5 correctly answered all 1,200 multiplication problems involving numbers up to 20 digits each without external tools. They said this required maximum reasoning. For two 100-digit numbers, they reported about 80% accuracy and suggested the drop likely came from exceeding a 128,000-token limit.

    The same user later reported roughly 0% accuracy for Fable 5.1 on 5-digit-by-6-digit multiplication at low reasoning. They attributed the failures to skipped reasoning, saying the model answered larger problems correctly when it began producing chain-of-thought reasoning, before accuracy declined again near 20-digit-by-20-digit multiplication.

    A reply suggested an external classifier might decide whether to reason, rather than the model itself. The tester countered with a quotation from Claude documentation saying the model decides whether and how much to think.

    Combined views

    93.5K

    Combined views

    93.5K

    18 Sources, first seen 21d ago

    1.1K likes
    21d ago
    first seen 21d ago
    18 Sources, first seen 21d ago
    1.1K likes66 comments329 saves100 reposts
    66 comments
    329 saves
    100 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    18 Sources

    Maksym Andriushchenko@maksym_andr@PMinervini @hadif4r yes, then the LLM can use external tools! it's a fully solved problem.21d
    Pasquale Minervini@PMinervini@maksym_andr @hadif4r yeah but that was also an option 2 years ago, right? LLMs (in their current iteration) still struggle with that21d
    Lisan al Gaib@scaling01I think Fable and Astra are not available without thinking because you could make inferences about their size and depth but also because they are more vulnerable to all sorts of attacks and since only Fable can read the hidden Fable CoT you don't want a weak non-CoT fable to break that21d
    Bharath Ramsundar@rbhar90@maksym_andr That's fascinating, thank you for answering. Decent evidence it's actually learning a multiplication algorithm...21d
    elie@eliebakouchRT @maksym_andr: 💥 The multiplication saga continues! It turns out, some frontier LLMs do struggle with simple multiplications. Unlike GPT…20d
    j⧉nus@repligate@maksym_andr > The model decides that it shouldn't produce any reasoning and outputs the (wrong) answer directly Afaik, it might not be “the model” deciding this, but some kind of external classifier20d

    18 Sources

    Maksym Andriushchenko@maksym_andr@PMinervini @hadif4r yes, then the LLM can use external tools! it's a fully solved problem.21d
    Pasquale Minervini@PMinervini@maksym_andr @hadif4r yeah but that was also an option 2 years ago, right? LLMs (in their current iteration) still struggle with that21d
    Lisan al Gaib@scaling01I think Fable and Astra are not available without thinking because you could make inferences about their size and depth but also because they are more vulnerable to all sorts of attacks and since only Fable can read the hidden Fable CoT you don't want a weak non-CoT fable to break that21d
    Bharath Ramsundar@rbhar90@maksym_andr That's fascinating, thank you for answering. Decent evidence it's actually learning a multiplication algorithm...21d
    elie@eliebakouchRT @maksym_andr: 💥 The multiplication saga continues! It turns out, some frontier LLMs do struggle with simple multiplications. Unlike GPT…20d
    j⧉nus@repligate@maksym_andr > The model decides that it shouldn't produce any reasoning and outputs the (wrong) answer directly Afaik, it might not be “the model” deciding this, but some kind of external classifier20d

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.