germans released a model that's actually not terrible
it's small, and still worse than Qwen3.5, but it's very comparable to Nemotron 3 Nano
but I guess that's what 27T pre-training tokens will do to a model https://twitter.com/effi288/status/2075904321707798699
For anyone wondering why they did not post-train this is actual sovereignty in the good sense of the world: slowly building strategic autonomy up to the pretraining level, while having the humility to start from the best practices in the open.…