• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Teknium Questions Astra Token Efficiency Versus Fable 5.1

    Nous Research engineer compares cache costs of two AI models in tweet.

    OK
    DP
    T🪽
    9 Sources, 26d ago, first seen 26d ago

    TLDR

    Teknium, co-founder and lead engineer at Nous Research and creator of the Hermes LLM family, posted on X asking users experienced with Fable 5.1 and Astra whether Astra delivers 8x or better token efficiency. He notes Astra carries 8x higher cache read costs than Fable 5.1 and states the pricing makes him reluctant to use it. The post presents his direct observation and question to the community without additional confirmation or response data.

    Combined views

    286.9K

    9 Sources, first seen 26d ago

    Combined views

    286.9K

    9 Sources, first seen 26d ago

    2.2K likes
    2.2K likes
    221 comments
    257 saves
    26 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    221 comments
    257 saves
    26 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    9 Sources

    @TekniumFor those of you who've used Fable 5.1 and Astra - is Astra 8x or better token efficiency?? It's 8x more expensive cache read to fable 5.1 has me pretty scared to touch it
    @krishnanrohit@Teknium Better. But not 8x better.
    @jxnlco@Teknium 60% of fable costs per task on evals
    @sandersted@gabriel1 depends on the task, and whether imperfections/failures lead to retries
    @vikhyatkwith sol xhigh i was burning through weekly limit in ~1 day. with astra medium it's ~12 hours not sure if it's worth it, will know in a couple days. but it seems like i'm going to have to 2x my codex accounts
    @DimitrisPapailI maintain that F5.1 in CC is near an order of magnitude better at token usage than Astra. My Codex usage runs out in a day vs a week for CC. Codex is still very token wasteful and checks constantly for input even if you tell it "go to sleep and check once in X minutes"
    @lateinteraction@thsottiaux @TJeparskis that seems to ignore the length of your context? many of us love the ability of Codex to sustain very long-running (read: multi-month) chats, or otherwise work such that the median context window has well over 100k tokens here, the cost of even Astra med >> Sol xhigh, right?

    9 Sources

    @TekniumFor those of you who've used Fable 5.1 and Astra - is Astra 8x or better token efficiency?? It's 8x more expensive cache read to fable 5.1 has me pretty scared to touch it
    @krishnanrohit@Teknium Better. But not 8x better.
    @jxnlco@Teknium 60% of fable costs per task on evals
    @sandersted@gabriel1 depends on the task, and whether imperfections/failures lead to retries
    @vikhyatkwith sol xhigh i was burning through weekly limit in ~1 day. with astra medium it's ~12 hours not sure if it's worth it, will know in a couple days. but it seems like i'm going to have to 2x my codex accounts
    @DimitrisPapailI maintain that F5.1 in CC is near an order of magnitude better at token usage than Astra. My Codex usage runs out in a day vs a week for CC. Codex is still very token wasteful and checks constantly for input even if you tell it "go to sleep and check once in X minutes"
    @lateinteraction@thsottiaux @TJeparskis that seems to ignore the length of your context? many of us love the ability of Codex to sustain very long-running (read: multi-month) chats, or otherwise work such that the median context window has well over 100k tokens here, the cost of even Astra med >> Sol xhigh, right?