Mistral previews the 1T-parameter Large 4, Google open-sources multimodal EmbeddingGemma 2, Anthropic expands its Cyber Verification Program
Openness is moving at both ends today: open weights reach up to a 1T-parameter flagship and down to a sub-1B multimodal embedding model that runs on a phone, while on the closed side Anthropic uses tiered vetting to open stronger cyber capabilities to verified defenders.
A 1 trillion-parameter natively multimodal model with 49 billion active parameters, fluent in more than 160 languages including every official EU language.
Preview API pricing is $1.36 per million input tokens and $4.18 per million output tokens; Mistral says it will release the weights by the end of the month, after real-world red-teaming with security leaders and partners.
Coding: 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4; 59.9% on AutomationBench, which covers 657 business workflows.
Security: 82% on the AA Cyber Index test that asks a model to reproduce and then patch a real vulnerability, and 93% of Cybench's 40 challenges; trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs.
Builder's takeAt $1.36 in and $4.18 out, it's worth running PandaClaws' long-form generation and multilingual rewriting against it. But this is a preview with no weights or license yet, so I'd benchmark it on my own eval set through the API now and only weigh self-hosting once the weights and license terms land at month end.
Load only what you need: 270M for text and code, 440M with vision, 570M with audio, 740M for all modalities; the text and code backbone has an 8,192-token context window.
Outputs 768-dimensional vectors that can be truncated to 512, 256 or 128 via Matryoshka Representation Learning; 14% higher than EmbeddingGemma 1 on MTEB (Code) while keeping its multilingual text accuracy.
On device, text-only weights use about 191 MB of active RAM and the full multimodal model about 567 MB on a Pixel 11 Pro, with INT4 and INT8 weights from quantization-aware training.
Released under Apache 2.0 with weights on Hugging Face, and supported in Ollama, vLLM, Transformers, sentence-transformers, MLX, LiteRT and more.
Builder's takeThis one maps straight onto my AI Cloud Drive: with images, video, recordings and documents in one vector space, a single query can search across file types without a separate model per format. I'd swap the 270M text version in against my current embedder first, then load the vision module as needed; truncating to 128 dimensions cuts index storage, but measure the recall loss on your own retrieval set first.
Defense Access covers SOC work, incident response, malware reverse engineering and vulnerability analysis; companies, universities, government security teams, open-source maintainers and individual researchers with a disclosure record can apply, with responses targeted within a few days.
Red Team Access adds authorized penetration testing and red-teaming, for organizations only, with reviews taking a few weeks; Specialized Access covers safety-critical systems such as flight software, power grids and telecom, reviewed case by case with the US government.
On CyScenarioBench, the Defense tier blocked 46 of 50 tasks on Opus 5.5; the Red Team tier blocked none and completed 34 of 50, the same as with no safeguards.
Enrolled organizations must retain data for misuse monitoring; once Enterprise Frontier Safeguards arrives this fall, they can store it in cloud infrastructure they control. CVP is available on the Claude Platform, Vertex AI and Microsoft Foundry.
Builder's takeIf your team keeps getting blocked when using Claude for code audits or security triage, this is the proper route: apply for the Defense tier first, since it's reviewed fastest. Note that enrolling means data retention, so for products handling private user files like my AI Cloud Drive, run security testing against a scrubbed copy rather than real user data.
Researched and drafted with AI assistance; editorial standards and views set by Darius.