<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Open models · AI Brief · darius.wiki</title><link>https://www.darius.wiki/en/topics/open-models/</link><atom:link href="https://www.darius.wiki/en/topics/open-models/rss.xml" rel="self" type="application/rss+xml"/><description>Open weights, licenses and small models you can run locally</description><language>en</language><lastBuildDate>Fri, 09 Oct 2026 00:00:00 GMT</lastBuildDate><item><title>LightOn open-sources LightOnOCR-3 in 0.8B, 1B and 4B sizes under Apache 2.0, scoring up to 86.3 on olmOCR-Bench</title><link>https://www.darius.wiki/en/daily/2026-10-09/lightonocr-3/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-09/lightonocr-3/</guid><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><category>Open source</category><description>Small document-parsing models with a new grounding mode that returns coordinates, suited to turning PDFs and scans into structured text. Take: The 0.8B is less than a point behind the 4B, which makes it attractive for document parsing in AI Cloud Drive: I'll run a comparison on the scans and table-heavy PDFs users upload most, to see whether some parsing can move from a cloud API onto our own machines.</description></item><item><title>Liquid AI releases open d1 decision models: d1-3B answers in one forward pass, 8 ms per question on an RTX 4090</title><link>https://www.darius.wiki/en/daily/2026-10-09/liquid-ai-d1/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-09/liquid-ai-d1/</guid><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><category>Open source</category><description>Small models that answer structured questions instead of generating text, aimed at devices and the edge. Take: Yes-or-no calls like moderation, routing or whether to hand off to a human don't need a generative model every time. I'll try d1-3B on PandaClaws' pre-publish compliance checks, but only as an evaluation until the license is clear.</description></item><item><title>NVIDIA fine-tunes Nemotron to gold-level IOI 2026 (535.4/600) and IMO 2026 (30/42) results, releasing weights and training data</title><link>https://www.darius.wiki/en/daily/2026-10-09/nvidia-nemotron-ioi-imo/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-09/nvidia-nemotron-ioi-imo/</guid><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><category>Research</category><description>One open model family clears the gold bar in both programming and math olympiads, with the data and evaluation pipelines released too. Take: Most teams won't run a 550B model; the part worth taking is the method: on IOI 2025 the Nano model (30B total, 3B active) went from 280 after SFT to 468 with test-time strategies like GenCorrect. If you build coding products, try a generate-then-self-correct layer on your own eval set first.</description></item><item><title>Mistral Large 4 enters public preview: 1T total, 49B active, $1.36 per million input tokens, weights by month end</title><link>https://www.darius.wiki/en/daily/2026-10-07/mistral-large-4-preview/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-07/mistral-large-4-preview/</guid><pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate><category>Model release</category><description>Mistral's largest model to date, natively multimodal; the preview API is live on Mistral Studio and the weights are due by the end of the month. Take: At $1.36 in and $4.18 out, it's worth running PandaClaws' long-form generation and multilingual rewriting against it. But this is a preview with no weights or license yet, so I'd benchmark it on my own eval set through the API now and only weigh self-hosting once the weights and license terms land at month end.</description></item><item><title>Google open-sources EmbeddingGemma 2: a 270M-to-740M multimodal embedding model, 14% better on code retrieval, Apache 2.0</title><link>https://www.darius.wiki/en/daily/2026-10-07/embeddinggemma-2/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-07/embeddinggemma-2/</guid><pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate><category>Open source</category><description>A small embedding model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space, bringing multimodal semantic search on device. Take: This one maps straight onto my AI Cloud Drive: with images, video, recordings and documents in one vector space, a single query can search across file types without a separate model per format. I'd swap the 270M text version in against my current embedder first, then load the vision module as needed; truncating to 128 dimensions cuts index storage, but measure the recall loss on your own retrieval set first.</description></item><item><title>Aleph Alpha open-sources Kolibri: a 78B MoE with 3.46B active per token, validated to 1M-token context, Apache 2.0</title><link>https://www.darius.wiki/en/daily/2026-10-06/aleph-alpha-kolibri-1/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-06/aleph-alpha-kolibri-1/</guid><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><category>Open source</category><description>An open German-and-English model with a reasoning mode and tool calling; its FP8 weights take about 78 GB and serve on a single H200 or B200. Take: About 3B active per token, one GPU, Apache 2.0: this is the shape of model that suits privately deployed long-document Q&amp;A like my AI Cloud Drive. But it only covers German and English, so don't drop it into Chinese workloads; what I want from it is a real throughput number for sliding-window attention plus long context on my own hardware before choosing the next self-hosted model.</description></item><item><title>Ai2 open-sources AstaBrief 8B, which writes a cited research report in one pass, about 3.5x faster than its Claude-powered mode</title><link>https://www.darius.wiki/en/daily/2026-10-05/allenai-astabrief-8b/</link><guid isPermaLink="true">https://www.darius.wiki/en/daily/2026-10-05/allenai-astabrief-8b/</guid><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><category>Open source</category><description>An 8B open-weights model turns retrieved excerpts into a cited long-form report, with weights and training data released together. Take: Retrieve-then-write-a-cited-answer is the most common request in my AI Cloud Drive's document Q&amp;A. An 8B model that does it in one call means private deployment and a much lower bill. I'll compare citation accuracy on my own documents before replacing the large-model call.</description></item></channel></rss>