Cloudflare open-sources Clef-omni, a multimodal decision model, and cuts Clef-flash input from $0.09 to $0.038 per million tokens
Decision models, which score inputs against given options instead of generating text, can now take images, audio and video directly.
// Key points
- Clef models score inputs against schema-defined options rather than generating text. Clef-omni takes text, images, audio (wav, mp3) and video (mp4, webm) in one call; it is built on Qwen3-Omni-30B-A3B-Instruct with the speech output components removed, and its weights are on Hugging Face.
- On Workers AI, Clef-omni costs $0.15 per million input tokens; median latency is about 130 ms for text, about 150 ms for images, and about 1.5 seconds for a 21-second video with sound.
- Clef-flash drops to $0.038 per million input tokens, with the hosted context cut from 64k to 24k. Clef stays at $0.24 but its median latency is 1.7x to 2.0x faster, and its weights ship with SGLang launch commands for self-hosting.
Builder's takeRead alongside Liquid AI’s d1 from 10-09, decision models are becoming a category of their own: moderation, classification and routing only need one answer picked, not a generative model. I’d run a comparison with Clef-omni for image and video moderation in PandaClaws; anyone on Clef-flash should note the context is now 24k.
// Background · from #Open models
Full timeline →- Oct 10 Google open-sources ML Drift, its on-device GPU inference engine, under Apache 2.0, with up to 12% lower memory overhead for Gemma
- Oct 9 LightOn open-sources LightOnOCR-3 in 0.8B, 1B and 4B sizes under Apache 2.0, scoring up to 86.3 on olmOCR-Bench
- Oct 9 Liquid AI releases open d1 decision models: d1-3B answers in one forward pass, 8 ms per question on an RTX 4090