• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Fable 5.1 reportedly triggers bio safeguards when counting letters in random words

    A user says the model refused a letter-counting task involving a completely random sequence of words, calling Fable 5.1 “basically unusable” and hoping for a fix.

    Lisan al GaibLA
    Maksym AndriushchenkoMA
    5 Sources, 20d ago, first seen 20d ago

    TLDR

    A user reports that counting letters in a completely random sequence of words triggered Fable 5.1’s bio safeguards, prompting refusals. They hope the issue gets fixed. In a related text-passage comparison, the same user also argues that Astra used fewer tokens more effectively than Fable.

    Combined views

    6.5K

    5 Sources, first seen 20d ago

    Combined views

    6.5K

    5 Sources, first seen 20d ago

    120 likes
    120 likes
    11 comments
    25 saves
    5 reposts
    11 comments
    25 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    Maksym Andriushchenko@maksym_andrCan frontier LLMs count letters in text passages? This task is very simple but is out-of-distribution for frontier LLMs that are benchmaxxed on SWE-Bench and TerminalBench-like tasks, but not on simple counting! It turns out there is a large gap between the behavior of GPT-6-Astra and Fable 5.1 at low reasoning effort: Astra maintains very high accuracy (93% at 1280 characters), while Fable fails much more often (53% at 1280 characters). Interestingly, the Fable curve is not monotonic. This is again a failure of Fable's adaptive thinking: roughly until 320 character the model decides to not use any thinking tokens at all, while that would've increased the counting accuracy. The motivation for this experiment comes from understanding the no/low-CoT performance and whether avoiding using CoT would be sufficient for these models to implement complex solutions (e.g., scheming) without verbalizing this in the CoT. 1/520d
    Lisan al Gaib@scaling01@maksym_andr at this point you need to ask for access to the non-thinking versions of both models20d

    5 Sources

    Maksym Andriushchenko@maksym_andrCan frontier LLMs count letters in text passages? This task is very simple but is out-of-distribution for frontier LLMs that are benchmaxxed on SWE-Bench and TerminalBench-like tasks, but not on simple counting! It turns out there is a large gap between the behavior of GPT-6-Astra and Fable 5.1 at low reasoning effort: Astra maintains very high accuracy (93% at 1280 characters), while Fable fails much more often (53% at 1280 characters). Interestingly, the Fable curve is not monotonic. This is again a failure of Fable's adaptive thinking: roughly until 320 character the model decides to not use any thinking tokens at all, while that would've increased the counting accuracy. The motivation for this experiment comes from understanding the no/low-CoT performance and whether avoiding using CoT would be sufficient for these models to implement complex solutions (e.g., scheming) without verbalizing this in the CoT. 1/520d
    Lisan al Gaib@scaling01@maksym_andr at this point you need to ask for access to the non-thinking versions of both models20d