Open source #Open models Oct 10, 2026 · Originally published Oct 8

Google open-sources ML Drift, its on-device GPU inference engine, under Apache 2.0, with up to 12% lower memory overhead for Gemma

The GPU engine behind LiteRT is now open on its own, covering phones, the web and desktop.

Primary source Google Developers Blog · developers.googleblog.com Read the original ↗

// Key points

  • ML Drift is the core GPU acceleration engine in LiteRT and is also available as a standalone library, under Apache 2.0 on GitHub.
  • It supports OpenGL ES, OpenCL, Metal and WebGPU, covering Android and iOS, plus Windows and Linux via WebGPU and macOS via Metal (desktop in preview).
  • Stage-aware optimizations switch kernels between prefill and decode for edge LLMs; Gemma benchmarks show up to 12% lower memory overhead than other frameworks, and YouTube Shorts saw up to 40% lower average frame latency.
  • The legacy TFLite GPU delegate gets no new features, and the LiteRT accelerator is fully backward compatible with existing models.

Builder's takePaired with EmbeddingGemma 2 from 10-07, on-device semantic search now has both the model and the engine. Products like AI Cloud Drive can trial local embedding on the web via WebGPU and keep privacy-sensitive files on the user's device; projects still on the TFLite GPU delegate should schedule a migration.

// Background · from #Open models

Full timeline →
  1. Oct 9 LightOn open-sources LightOnOCR-3 in 0.8B, 1B and 4B sizes under Apache 2.0, scoring up to 86.3 on olmOCR-Bench
  2. Oct 9 Liquid AI releases open d1 decision models: d1-3B answers in one forward pass, 8 ms per question on an RTX 4090
  3. Oct 9 NVIDIA fine-tunes Nemotron to gold-level IOI 2026 (535.4/600) and IMO 2026 (30/42) results, releasing weights and training data