Aleph Alpha open-sources Kolibri: a 78B MoE with 3.46B active per token, validated to 1M-token context, Apache 2.0
An open German-and-English model with a reasoning mode and tool calling; its FP8 weights take about 78 GB and serve on a single H200 or B200.
// Key points
- Mixture-of-experts with 78B total parameters and 3.46B active per token; German and English only, a deliberate choice of depth over breadth, with a tokenizer built for German.
- Native context is 262,144 tokens; quality and serving efficiency are validated up to 1,048,576 tokens, with ≤262,144 recommended for latency-sensitive or complex tasks.
- FP8 weights need about 78 GB: minimum 2× A100 80 GB, 2× H100, or one H200, B200 or B300. Served through vLLM with Aleph Alpha's plugin and an OpenAI-compatible API.
- Reasoning effort can be set to low, medium, high or off, and tool calling can be combined with it; pre-trained on 20T tokens, released October 3 under Apache 2.0.
Builder's takeAbout 3B active per token, one GPU, Apache 2.0: this is the shape of model that suits privately deployed long-document Q&A like my AI Cloud Drive. But it only covers German and English, so don't drop it into Chinese workloads; what I want from it is a real throughput number for sliding-window attention plus long context on my own hardware before choosing the next self-hosted model.