• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researcher Completes Exact Claude Tokenizer Reproduction

    ctok tool now matches Claude output for natural languages and programming languages.

    LB
    SZ
    YA
    19 Sources, 48d ago, first seen 48d ago

    TLDR

    Sander Land announced that ctok now exactly reproduces token counts from Claude's tokenizer. The match holds across 500+ natural languages and 22 programming languages for both tokenizer families. Vocabulary estimates for Claude 4.7+ dropped to 15k entries. The tool runs offline without API calls. Posts from other researchers and engineers on X note the small vocabulary size and discuss possible links to softmax computation, while some question whether code comments in the project were AI-generated.

    Combined views

    359.4K

    19 Sources, first seen 48d ago

    Combined views

    359.4K

    19 Sources, first seen 48d ago

    1.6K likes
    1.6K likes
    56 comments
    1.2K saves
    154 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    56 comments
    1.2K saves
    154 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    19 Sources

    @magikarp_tokens✅ Claude's tokenizer reproduction is done! ctok now exactly reproduces token counts across 500+ natural languages and 22 programming languages, for both tokenizer families. The vocabulary estimates even went down: just 15k entries for Claude 4.7+ 🤯 🔗 https://tokenize.rs/claude
    @andersonbcdefgRT @magikarp_tokens: ✅ Claude's tokenizer reproduction is done! ctok now exactly reproduces token counts across 500+ natural languages and…
    @yoavartziAlmost thinking that Anthro figured out @nthngdy's backprop bottleneck for a while now, and ran with a smaller vocab for better learning... https://arxiv.org/abs/2603.10145
    @nrehiew_It seems likely that Claude 4.7+ having such a small vocab size is Anthropic's attempt to deal with the softmax bottleneck. A fun fact is that the original softmax bottleneck paper was led by Yang Zhilin, the founder of Moonshot AI
    @samsja19@nrehiew_ hmmm in what way is going so low solving a bottelneck ? I don't think that's the reason
    @recurseparadox@nrehiew_ This goes away for larger models where the dmodel is several thousand
    @xlr8harderMost interesting part of this is what it implies about the tradeoffs they are balancing. Giant vocabs are painful to train efficiently, but lengthening docs in context is a major cost (though ameliorated by hybrid attn). We need someone to do this work in public.
    @suchenzanghmm... did folks read this and think: "oh yeah obviously, that makes perfect sense, file is self-describing, the anchors are not summed, they are subtracted as a template constant, README.md explains the public contract, yup sounds right to me!" or "eh, comments are all slop now anyways, it's more important that (we think) we've found claude's 15k vocab!"
    @xeophon@suchenzang Eh, I expect all code to be generated
    @giffmana@yoavartzi @nthngdy I think this is only a small model issue, for large model dim it goes away, even in your paper experiments no?

    19 Sources

    @magikarp_tokens✅ Claude's tokenizer reproduction is done! ctok now exactly reproduces token counts across 500+ natural languages and 22 programming languages, for both tokenizer families. The vocabulary estimates even went down: just 15k entries for Claude 4.7+ 🤯 🔗 https://tokenize.rs/claude
    @andersonbcdefgRT @magikarp_tokens: ✅ Claude's tokenizer reproduction is done! ctok now exactly reproduces token counts across 500+ natural languages and…
    @yoavartziAlmost thinking that Anthro figured out @nthngdy's backprop bottleneck for a while now, and ran with a smaller vocab for better learning... https://arxiv.org/abs/2603.10145
    @nrehiew_It seems likely that Claude 4.7+ having such a small vocab size is Anthropic's attempt to deal with the softmax bottleneck. A fun fact is that the original softmax bottleneck paper was led by Yang Zhilin, the founder of Moonshot AI
    @samsja19@nrehiew_ hmmm in what way is going so low solving a bottelneck ? I don't think that's the reason
    @recurseparadox@nrehiew_ This goes away for larger models where the dmodel is several thousand
    @xlr8harderMost interesting part of this is what it implies about the tradeoffs they are balancing. Giant vocabs are painful to train efficiently, but lengthening docs in context is a major cost (though ameliorated by hybrid attn). We need someone to do this work in public.
    @suchenzanghmm... did folks read this and think: "oh yeah obviously, that makes perfect sense, file is self-describing, the anchors are not summed, they are subtracted as a template constant, README.md explains the public contract, yup sounds right to me!" or "eh, comments are all slop now anyways, it's more important that (we think) we've found claude's 15k vocab!"
    @xeophon@suchenzang Eh, I expect all code to be generated
    @giffmana@yoavartzi @nthngdy I think this is only a small model issue, for large model dim it goes away, even in your paper experiments no?