Google open-sources ML Drift, its on-device GPU inference engine, under Apache 2.0, with up to 12% lower memory overhead for Gemma
The GPU engine behind LiteRT is now open on its own, covering phones, the web and desktop.
// Key points
- ML Drift is the core GPU acceleration engine in LiteRT and is also available as a standalone library, under Apache 2.0 on GitHub.
- It supports OpenGL ES, OpenCL, Metal and WebGPU, covering Android and iOS, plus Windows and Linux via WebGPU and macOS via Metal (desktop in preview).
- Stage-aware optimizations switch kernels between prefill and decode for edge LLMs; Gemma benchmarks show up to 12% lower memory overhead than other frameworks, and YouTube Shorts saw up to 40% lower average frame latency.
- The legacy TFLite GPU delegate gets no new features, and the LiteRT accelerator is fully backward compatible with existing models.
Builder's takePaired with EmbeddingGemma 2 from 10-07, on-device semantic search now has both the model and the engine. Products like AI Cloud Drive can trial local embedding on the web via WebGPU and keep privacy-sensitive files on the user's device; projects still on the TFLite GPU delegate should schedule a migration.
// Background · from #Open models
Full timeline →- Oct 9 LightOn open-sources LightOnOCR-3 in 0.8B, 1B and 4B sizes under Apache 2.0, scoring up to 86.3 on olmOCR-Bench
- Oct 9 Liquid AI releases open d1 decision models: d1-3B answers in one forward pass, 8 ms per question on an RTX 4090
- Oct 9 NVIDIA fine-tunes Nemotron to gold-level IOI 2026 (535.4/600) and IMO 2026 (30/42) results, releasing weights and training data