• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Aleph Alpha releases open-weight Kolibri model under Apache 2.0

    Aleph Alpha says the bilingual mixture-of-experts model has 78.1 billion parameters and a context option up to 1 million tokens.

    C🤗
    EM
    AG
    12 Sources, ,

    TLDR

    Aleph Alpha released Kolibri, an English-German open-weight model under Apache 2.0. The company says the mixture-of-experts system has 78.1 billion total parameters, with 3.46 billion active per token. Its technical table lists a 262,144-token longest trained length, with serving instructions for configurations up to 1 million. Kolibri is available through Hugging Face and positioned for on-premise enterprise and government use.

    Combined views

    635.9K

    12 Sources, first seen 13h ago

    Combined views

    635.9K

    12 Sources, first seen 13h ago

    3.7K likes
    13h ago
    first seen 13h ago
    3.7K likes
    206 comments
    1.9K saves
    523 reposts
    206 comments
    1.9K saves
    523 reposts

    Aleph Alpha released Kolibri on Oct. 3, offering the English-German model’s full weights on Hugging Face under Apache 2.0 license terms. The company is pitching the model to enterprise and government customers that want to run it on their own infrastructure.

    A sparse model with a long context option

    Kolibri is a mixture-of-experts model, an architecture that activates only part of the network for each token. According to Aleph Alpha’s launch article, it has 78.1 billion total parameters but activates 3.46 billion per token across six of its 384 experts.

    Featured Source

    The company advertises support for contexts up to 1 million tokens. Its technical table lists 262,144 tokens as the longest trained length, and the serving instructions require an extra configuration flag for contexts beyond that point. The same table lists four reasoning settings: none, low, medium and high.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #2

    Today's Rank

    #2

    Aleph Alpha says it trained Kolibri on 768 B200 GPUs in three stages. Those included 20 trillion tokens of pretraining at a 16,384-token sequence length, 3.44 trillion tokens of mid-training at 65,536 tokens and 200 billion tokens of long-context adaptation at 262,144 tokens.

    Built around local control

    Aleph Alpha says its teams built the model in Germany and trained it on infrastructure in Germany and Finland. It presents that development chain, along with support for on-premise deployment, as part of Kolibri’s “sovereign” positioning for regulated industries.

    Running Kolibri requires Aleph Alpha’s inference package and a compatible vLLM setup. The company provides a container image, a Python package and serving commands for reasoning and tool use, giving organizations a path to operate the model without sending internal data to a third-party inference service.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    13 Sources

    Aleph AlphaKolibri Has Landed: A Sovereign Open-Weight Model — Aleph Alpha
    @Aleph__AlphaSmall bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0.13h
    @ClementDelangueRT @Aleph__Alpha: Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now…10h
    @antgrassoYou found a new AI model with billions of parameters and you like what you see, but now you are wondering whether it is also efficient. Before getting excited about the size, ask one more question: what does it take to run that model at the scale you need? Microblog @antgrasso That number tells you only part of what matters. Parameters are numerical values learned during training. More of them can increase what a model is capable of doing, but they also have to be stored and used when the model runs. That requires memory and computation. How long does the model take to produce an answer? How much energy does it use? What does running it cost at the scale you need? Efficiency also depends on the workload. A model can be a sensible choice for one task and unnecessarily demanding for another. The same model can behave differently depending on the hardware and how it is deployed. So parameter count can tell you something about model size and capacity. It cannot tell you, by itself, whether the model is efficient for the job you need to do. Big numbers may impress. Efficiency is about the resources required to get the result you need. #AIEfficiency9h
    @aidangomezCongrats to the AlephAlpha team on a very cool German-English model and very detailed technical report! Check it out and give the model a spin!7h
    @teortaxesTexHybrid SWA, 3.46B active. Europe on the Pareto frontier indeed. At least they're keeping up with the Chinese literature, kind of (but Sberbank is on MLA + GDN already; Euros have too much GDP to stoop that low). Scale that up.7h
    @nickfrosstNew Apache 2.0 model from aleph alpha. It rocks. This is the way. So excited to build together.5h
    @testingcatalogAleph Alpha released Kolibri, a 78B-parameter, 3.46B active, MoE open-weight model with 1M context window, distributed under the Apache 2.0 license. Aleph Alpha is a German AI lab, and it is probably the first general-purpose model from the EU, apart from Mistral, to achieve this level of performance! > "The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency." > "The Aleph Alpha team developed a bilingual German/English tokenizer and focused on including organic German data throughout the training process of the model, so that 21.3% of the pre-training tokens are German." > "Kolibri was built with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind from the ground up, with copyright law being a focus of our work to trustworthy technology." More EU labs are joining the party 👀3h
    @cohereNew Apache 2.0 model from @Aleph__Alpha. It's really good. If you like that, you're going to love what we build together.3h
    @AICoffeeBreakKolibri is out 🐦 (Apache 2.0) and new research ideas around measuring grounding and reducing hallucinations with Merlin-Arthur made it into the model. This is just the beginning: lots of open questions, lots still to improve. But getting research ideas into an actual model, learning from them, and having an amazing in-house model training pipeline to iterate quickly is the fun part.2h

    13 Sources

    Aleph AlphaKolibri Has Landed: A Sovereign Open-Weight Model — Aleph Alpha
    @Aleph__AlphaSmall bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0.13h
    @ClementDelangueRT @Aleph__Alpha: Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now…10h
    @antgrassoYou found a new AI model with billions of parameters and you like what you see, but now you are wondering whether it is also efficient. Before getting excited about the size, ask one more question: what does it take to run that model at the scale you need? Microblog @antgrasso That number tells you only part of what matters. Parameters are numerical values learned during training. More of them can increase what a model is capable of doing, but they also have to be stored and used when the model runs. That requires memory and computation. How long does the model take to produce an answer? How much energy does it use? What does running it cost at the scale you need? Efficiency also depends on the workload. A model can be a sensible choice for one task and unnecessarily demanding for another. The same model can behave differently depending on the hardware and how it is deployed. So parameter count can tell you something about model size and capacity. It cannot tell you, by itself, whether the model is efficient for the job you need to do. Big numbers may impress. Efficiency is about the resources required to get the result you need. #AIEfficiency9h
    @aidangomezCongrats to the AlephAlpha team on a very cool German-English model and very detailed technical report! Check it out and give the model a spin!7h
    @teortaxesTexHybrid SWA, 3.46B active. Europe on the Pareto frontier indeed. At least they're keeping up with the Chinese literature, kind of (but Sberbank is on MLA + GDN already; Euros have too much GDP to stoop that low). Scale that up.7h
    @nickfrosstNew Apache 2.0 model from aleph alpha. It rocks. This is the way. So excited to build together.5h
    @testingcatalogAleph Alpha released Kolibri, a 78B-parameter, 3.46B active, MoE open-weight model with 1M context window, distributed under the Apache 2.0 license. Aleph Alpha is a German AI lab, and it is probably the first general-purpose model from the EU, apart from Mistral, to achieve this level of performance! > "The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency." > "The Aleph Alpha team developed a bilingual German/English tokenizer and focused on including organic German data throughout the training process of the model, so that 21.3% of the pre-training tokens are German." > "Kolibri was built with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind from the ground up, with copyright law being a focus of our work to trustworthy technology." More EU labs are joining the party 👀3h
    @cohereNew Apache 2.0 model from @Aleph__Alpha. It's really good. If you like that, you're going to love what we build together.3h
    @AICoffeeBreakKolibri is out 🐦 (Apache 2.0) and new research ideas around measuring grounding and reducing hallucinations with Merlin-Arthur made it into the model. This is just the beginning: lots of open questions, lots still to improve. But getting research ideas into an actual model, learning from them, and having an amazing in-house model training pipeline to iterate quickly is the fun part.2h