Open source #Open models Oct 6, 2026 · Originally published Oct 3

Aleph Alpha open-sources Kolibri: a 78B MoE with 3.46B active per token, validated to 1M-token context, Apache 2.0

An open German-and-English model with a reasoning mode and tool calling; its FP8 weights take about 78 GB and serve on a single H200 or B200.

Primary source Aleph Alpha on Hugging Face · huggingface.co Read the original ↗

// Key points

  • Mixture-of-experts with 78B total parameters and 3.46B active per token; German and English only, a deliberate choice of depth over breadth, with a tokenizer built for German.
  • Native context is 262,144 tokens; quality and serving efficiency are validated up to 1,048,576 tokens, with ≤262,144 recommended for latency-sensitive or complex tasks.
  • FP8 weights need about 78 GB: minimum 2× A100 80 GB, 2× H100, or one H200, B200 or B300. Served through vLLM with Aleph Alpha's plugin and an OpenAI-compatible API.
  • Reasoning effort can be set to low, medium, high or off, and tool calling can be combined with it; pre-trained on 20T tokens, released October 3 under Apache 2.0.

Builder's takeAbout 3B active per token, one GPU, Apache 2.0: this is the shape of model that suits privately deployed long-document Q&A like my AI Cloud Drive. But it only covers German and English, so don't drop it into Chinese workloads; what I want from it is a real throughput number for sliding-window attention plus long context on my own hardware before choosing the next self-hosted model.

// Background · from #Open models

Full timeline →
  1. Oct 5 Ai2 open-sources AstaBrief 8B, which writes a cited research report in one pass, about 3.5x faster than its Claude-powered mode