Open models
Open weights, licenses and small models you can run locally. Every brief item about Open models, newest first, each with its primary source.
Where it stands
Open models kept moving at both ends this week. At the top, Mistral Large 4 entered public preview at around a trillion total parameters with weights due at month end, and NVIDIA's Nemotron (Ultra-CC, 550B total and 55B active) scored 535.4/600 at IOI 2026 and 30/42 at IMO 2026, with weights and training data released. At the bottom, October 7 and 8 brought several specialized small models: LightOnOCR-3 (0.8B to 4B, Apache 2.0, up to 86.3 on olmOCR-Bench) and Liquid AI's d1 decision models (d1-3B answers in 8 ms on an RTX 4090), on top of EmbeddingGemma 2 and Kolibri earlier.
Small, specialized models are becoming genuinely useful: OCR, embeddings and yes-or-no decisions each have dedicated models under 4B that approach or beat much larger models on their task. Licenses are still uneven, though: LightOnOCR-3 is Apache 2.0, while the d1 and Nemotron posts don't state a license.
What to watch next: the license and self-hosting requirements for Large 4's weights at month end, the licenses on d1 and Nemotron weights, and whether small models hold up on real business data the way they do on leaderboards.
My advice: list your fixed tasks like OCR, classification, moderation and retrieval, and swap-test each with a specialized model under 4B, confirming the license before production; keep evaluating flagship open models as swappable API vendors until weights and licenses land.
Darius · Updated Oct 9, 2026
Timeline 7 items
-
Oct 9 · Fri · 3 items
- Open source huggingface.co ↗LightOn open-sources LightOnOCR-3 in 0.8B, 1B and 4B sizes under Apache 2.0, scoring up to 86.3 on olmOCR-Bench
Take · The 0.8B is less than a point behind the 4B, which makes it attractive for document parsing in AI Cloud Drive: I'll run a comparison on the scans and table-heavy PDFs users upload most, to see whether some parsing can move from a cloud API onto our own machines.
From the Oct 9 brief · item 06 → - Open source huggingface.co ↗Liquid AI releases open d1 decision models: d1-3B answers in one forward pass, 8 ms per question on an RTX 4090
Take · Yes-or-no calls like moderation, routing or whether to hand off to a human don't need a generative model every time. I'll try d1-3B on PandaClaws' pre-publish compliance checks, but only as an evaluation until the license is clear.
From the Oct 9 brief · item 07 → - Research huggingface.co ↗NVIDIA fine-tunes Nemotron to gold-level IOI 2026 (535.4/600) and IMO 2026 (30/42) results, releasing weights and training data
Take · Most teams won't run a 550B model; the part worth taking is the method: on IOI 2025 the Nano model (30B total, 3B active) went from 280 after SFT to 468 with test-time strategies like GenCorrect. If you build coding products, try a generate-then-self-correct layer on your own eval set first.
From the Oct 9 brief · item 08 →
-
-
Oct 7 · Wed · 2 items
- Model release mistral.ai ↗Mistral Large 4 enters public preview: 1T total, 49B active, $1.36 per million input tokens, weights by month end
Take · At $1.36 in and $4.18 out, it's worth running PandaClaws' long-form generation and multilingual rewriting against it. But this is a preview with no weights or license yet, so I'd benchmark it on my own eval set through the API now and only weigh self-hosting once the weights and license terms land at month end.
From the Oct 7 brief · item 01 → - Open source developers.googleblog.com ↗Google open-sources EmbeddingGemma 2: a 270M-to-740M multimodal embedding model, 14% better on code retrieval, Apache 2.0
Take · This one maps straight onto my AI Cloud Drive: with images, video, recordings and documents in one vector space, a single query can search across file types without a separate model per format. I'd swap the 270M text version in against my current embedder first, then load the vision module as needed; truncating to 128 dimensions cuts index storage, but measure the recall loss on your own retrieval set first.
From the Oct 7 brief · item 02 →
-
-
Oct 6 · Tue · 1 item
- Open source huggingface.co ↗Aleph Alpha open-sources Kolibri: a 78B MoE with 3.46B active per token, validated to 1M-token context, Apache 2.0
Take · About 3B active per token, one GPU, Apache 2.0: this is the shape of model that suits privately deployed long-document Q&A like my AI Cloud Drive. But it only covers German and English, so don't drop it into Chinese workloads; what I want from it is a real throughput number for sliding-window attention plus long context on my own hardware before choosing the next self-hosted model.
From the Oct 6 brief · item 01 →
-
-
Oct 5 · Mon · 1 item
- Open source huggingface.co ↗Ai2 open-sources AstaBrief 8B, which writes a cited research report in one pass, about 3.5x faster than its Claude-powered mode
Take · Retrieve-then-write-a-cited-answer is the most common request in my AI Cloud Drive's document Q&A. An 8B model that does it in one call means private deployment and a much lower bill. I'll compare citation accuracy on my own documents before replacing the large-model call.
From the Oct 5 brief · item 02 →
-